Semantic Structure Document Similarity Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current document processing systems are limited in comparing documents across different languages and formats, failing to accurately detect similarity or difference, especially in cross-language plagiarism detection and handling diverse types of information within documents.

Innovation Solution

A method that estimates and visualizes similarity/difference between documents using syntactic and semantic analyses, converting textual parts into language-independent semantic structures, and comparing various types of information such as text, images, and video, allowing for cross-language and cross-format document comparison without requiring machine translation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If lexical-based similarity measures are used for document comparison, then the computation is simple and fast, but the accuracy fails for cross-language documents and documents with different formats

Engineering Contradiction:
Improvedocument similarity detection accuracyVSAvoidcomparison system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary representation layer called 'semantic structure' that mediates between raw lexical content and similarity comparison. This semantic structure serves as a universal intermediary that can represent documents in different languages and formats in a unified manner, enabling accurate cross-language and cross-format similarity detection without requiring complex language-specific or format-specific processing for each comparison task.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a universal document representation framework that can handle multiple languages and formats through a single semantic structure paradigm. This universal approach allows the same comparison mechanism to work across diverse document types (text, images, video, audio) and languages without requiring separate specialized systems for each, thereby improving accuracy while controlling complexity through consolidation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If machine translation is used for cross-language document comparison, then language barriers can be overcome, but translation errors propagate and reduce similarity detection accuracy

Engineering Contradiction:
Improvecross-language comparison capabilityVSAvoidsimilarity detection accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent replaces the mechanical translation process with a direct semantic structure comparison approach. Instead of translating documents and then comparing them (which propagates translation errors), the system extracts semantic structures directly from source documents in their original languages and compares these structures. This substitution eliminates the translation step entirely while maintaining cross-language comparison capability, thereby avoiding translation error propagation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If comprehensive syntactic and semantic analysis is performed on documents, then the similarity estimation accuracy is improved, but the processing time increases significantly

Engineering Contradiction:
Improvesimilarity estimation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary extraction and representation of semantic structures for all documents in advance, creating a standardized semantic representation format. This preliminary action transforms the complex task of accurate similarity comparison into a more efficient process, as the heavy syntactic and semantic analysis is done upfront during document indexing, while actual similarity queries can then proceed faster by comparing pre-computed semantic structures rather than performing analysis in real-time.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If documents are compared based on lexical features only, then the processing is fast and simple, but nuanced differences in meaning and intent are missed

Engineering Contradiction:
Improveprocessing speedVSAvoidmeaning-based similarity detection
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent adds another dimension to document representation by introducing semantic structure as a higher-level abstraction layer above lexical features. This dimensional change allows the system to operate simultaneously at multiple levels: lexical level for fast initial filtering and semantic structure level for precise meaning-based comparison. The semantic structure dimension captures nuanced differences in meaning, intent, and context that lexical features alone cannot detect, while maintaining processing efficiency through hierarchical organization.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS9189482B2Similar document search
Publication Date: 2015.11.17 ABBYY DEVELOPMENT INC
  • US9189482B2 patent drawing
  • US9189482B2 patent drawing
  • US9189482B2 patent drawing

AI summary

Described herein are methods for finding substantially similar/different sources (files and documents), and estimating similarity or difference between given sources. Similarity and difference may be found across a variety of formats. Sources may be in one or more languages such that similarity and difference may be found across any number and types of languages. A variety of characteristics may be used to arrive at an overall measure of similarity or difference including determining or identifying syntactic roles, semantic roles and semantic classes in reference to sources.