Text Fragment Similarity Detection for Document Reuse
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text management systems do not efficiently facilitate the reuse of text content by comparing text fragments in real-time during document creation, limiting the potential for reusing text stored in databases.
Innovation Solution
A method and system that compares text fragments entered by an author to previously stored text fragments using edit distance and word occurrence algorithms, presenting similar alternatives for substitution, and allowing optional substitution based on user choice, with the system capable of storing and indexing text fragments for fast similarity searches.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If text management systems use traditional plagiarism detection methods to compare documents, then they can detect duplicate paragraphs, but they cannot efficiently facilitate real-time text re-use during document creation
Solution Approach 1:
The system pre-processes and stores text fragments from existing documents in a database before they are needed. This preliminary indexing and fragmentation of text allows for rapid retrieval and comparison during document creation, eliminating the need for authors to manually research existing text storage databases.
Solution Approach 2:
The invention replaces manual text searching and comparison with automated computational algorithms including edit distance calculations, word occurrence analysis, and word difference algorithms. This substitution of mechanical author research with automated text fragment comparison systems dramatically improves text re-use efficiency.
2Measurement precision
If the system compares text fragments using multiple algorithms (edit distance, word occurrence, word difference), then similarity detection accuracy improves, but system complexity increases
Solution Approach 1:
The text comparison function is segmented into three distinct algorithmic components: edit distance calculation, word occurrence analysis, and word difference analysis. Each algorithm handles a specific aspect of similarity detection, allowing them to be developed, optimized, and maintained independently while working together to provide comprehensive text comparison.
Solution Approach 2:
The system employs multiple comparison algorithms that serve universal text similarity detection needs. Each algorithm (edit distance, word occurrence, word difference) can independently evaluate different aspects of text similarity, and their results can be combined or used separately depending on the specific comparison requirements, providing multi-functional capability.
Data Source
AI summary
Text management software system which stores and retrieves text as paragraph size Text Fragments. Each fragment is stored in a separate record in a text store. As text is created and is to be saved to a text store, the system is adapted to break the text into paragraph size text fragments. As or before text is stored the system runs statistical comparisons between text fragments and builds a matrix of similarity between the text fragments stored in the separate records and presents these to the author for substitution.


