Document Comparison Using Unique Roots and Periodic Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for comparing documents with different formatting and structures, such as Word and PDF, are inefficient and fail to accurately identify differences, especially in compound documents with repetitive base subdocuments, leading to poor matching and understanding of text flow.

Innovation Solution

A method that systematically compares documents by identifying unique roots, expanding matches around them, and using histograms to determine the repetition period of base subdocuments, allowing for accurate alignment and visualization of agreements and differences through lists, and optimizing the comparison algorithm based on the repetition period.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional text-comparing algorithms monitor document flow linearly, then the comparison process is simple, but the accuracy of detecting text differences deteriorates when documents have different formatting or text arrangements

Engineering Contradiction:
Improvecomparison process simplicityVSAvoidtext difference detection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the document into text passages or blocks that can be independently identified and matched. Instead of treating the document as a continuous linear flow, it divides the content into discrete units (text passages) that can be located using unique roots and matched between documents regardless of their positional arrangement, thereby resolving the contradiction between simple processing and accurate detection.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary structure (the list of text passages with their locations and content) that mediates between the raw document text and the comparison process. This intermediary allows the system to handle documents with different formatting by normalizing the representation of text passages before comparison, improving detection accuracy without significantly complicating the overall process.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the algorithm uses unique roots to identify text passages, then the matching accuracy improves, but the complexity of handling compound documents with repetitive structures worsens

Engineering Contradiction:
Improvetext passage matching accuracyVSAvoidalgorithm complexity for repetitive structures
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies periodic action by detecting the repetition period n and systematically processing text passages at regular intervals. The algorithm identifies that certain text passages repeat every n occurrences and processes them in a periodic manner, which simplifies the handling of compound documents with repetitive structures while maintaining matching accuracy through the use of unique roots within each period.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The patent changes the parameter of root occurrence frequency from requiring exactly once (traditional approach) to allowing exactly n times (where n is the repetition period). This parameter change enables the algorithm to handle repetitive structures effectively by adjusting the uniqueness criterion to match the document's structural characteristics, thereby reducing complexity without sacrificing matching precision.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If the algorithm processes all text passages to ensure complete matching, then the comprehensiveness of comparison improves, but the time required for comparison increases

Engineering Contradiction:
Improvecomparison completenessVSAvoidcomparison processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-identifying and marking text passages using unique roots before the main comparison process. By locating and marking these key passages in advance, the algorithm can then efficiently compare only the marked passages and their surrounding context, rather than processing every single text passage, thus maintaining completeness while reducing processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by focusing the detailed comparison effort on text passages that contain or are near unique roots, rather than uniformly processing all passages. The expansion distance parameter allows the algorithm to extend the comparison scope partially beyond the roots themselves, ensuring completeness for relevant passages while avoiding unnecessary processing of unrelated text, thereby optimizing the time-completeness tradeoff.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10474672B2Method for comparing text files with differently arranged text sections in documents
Publication Date: 2019.11.12 SCHLAFENDER HASE GMBH SOFTWARE & COMM
  • US10474672B2 patent drawing
  • US10474672B2 patent drawing
  • US10474672B2 patent drawing

AI summary

A method for comparing and analyzing digital documents includes searching for unambiguous roots in both documents. These roots are unique units that occur in both documents. The roots can be individual words, word groups or other unambiguous textual formatting functions. There is then a search for identical roots in the other document (Root1 from Content1, and Root2 from Content2, with Root1=Root2). If a pair is found, the area around these roots is compared until there is no longer any agreement. During the area search, both preceding words and subsequent words are analyzed. The areas that are found in this way, Area1 around Root1 and Area2 around Root2, are stored in lists, List1 and List2, allocated to Doc1 and Doc2. This procedure is repeated until no roots can be found any longer. The result is either a remaining area that has no overlaps, or complete identity of the documents.