Document Comparison Using Unique Roots and Periodic Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for comparing documents with different formatting and structures, such as Word and PDF, are inefficient and fail to accurately identify differences, especially in compound documents with repetitive base subdocuments, leading to poor matching and understanding of text flow.
Innovation Solution
A method that systematically compares documents by identifying unique roots, expanding matches around them, and using histograms to determine the repetition period of base subdocuments, allowing for accurate alignment and visualization of agreements and differences through lists, and optimizing the comparison algorithm based on the repetition period.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional text-comparing algorithms monitor document flow linearly, then the comparison process is simple, but the accuracy of detecting text differences deteriorates when documents have different formatting or text arrangements
Solution Approach 1:
The patent segments the document into text passages or blocks that can be independently identified and matched. Instead of treating the document as a continuous linear flow, it divides the content into discrete units (text passages) that can be located using unique roots and matched between documents regardless of their positional arrangement, thereby resolving the contradiction between simple processing and accurate detection.
Solution Approach 2:
The patent introduces an intermediary structure (the list of text passages with their locations and content) that mediates between the raw document text and the comparison process. This intermediary allows the system to handle documents with different formatting by normalizing the representation of text passages before comparison, improving detection accuracy without significantly complicating the overall process.
2Measurement precision
If the algorithm uses unique roots to identify text passages, then the matching accuracy improves, but the complexity of handling compound documents with repetitive structures worsens
Solution Approach 1:
The patent applies periodic action by detecting the repetition period n and systematically processing text passages at regular intervals. The algorithm identifies that certain text passages repeat every n occurrences and processes them in a periodic manner, which simplifies the handling of compound documents with repetitive structures while maintaining matching accuracy through the use of unique roots within each period.
Solution Approach 2:
The patent changes the parameter of root occurrence frequency from requiring exactly once (traditional approach) to allowing exactly n times (where n is the repetition period). This parameter change enables the algorithm to handle repetitive structures effectively by adjusting the uniqueness criterion to match the document's structural characteristics, thereby reducing complexity without sacrificing matching precision.
3Reliability
If the algorithm processes all text passages to ensure complete matching, then the comprehensiveness of comparison improves, but the time required for comparison increases
Solution Approach 1:
The patent performs preliminary action by pre-identifying and marking text passages using unique roots before the main comparison process. By locating and marking these key passages in advance, the algorithm can then efficiently compare only the marked passages and their surrounding context, rather than processing every single text passage, thus maintaining completeness while reducing processing time.
Solution Approach 2:
The patent applies partial action by focusing the detailed comparison effort on text passages that contain or are near unique roots, rather than uniformly processing all passages. The expansion distance parameter allows the algorithm to extend the comparison scope partially beyond the roots themselves, ensuring completeness for relevant passages while avoiding unnecessary processing of unrelated text, thereby optimizing the time-completeness tradeoff.
Data Source
AI summary
A method for comparing and analyzing digital documents includes searching for unambiguous roots in both documents. These roots are unique units that occur in both documents. The roots can be individual words, word groups or other unambiguous textual formatting functions. There is then a search for identical roots in the other document (Root1 from Content1, and Root2 from Content2, with Root1=Root2). If a pair is found, the area around these roots is compared until there is no longer any agreement. During the area search, both preceding words and subsequent words are analyzed. The areas that are found in this way, Area1 around Root1 and Area2 around Root2, are stored in lists, List1 and List2, allocated to Doc1 and Doc2. This procedure is repeated until no roots can be found any longer. The result is either a remaining area that has no overlaps, or complete identity of the documents.


