Archive Storage Boundary Segmentation for Duplicate Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for identifying and reducing duplicate copies in archival storage systems are inefficient, leading to long delays and excessive resource occupation due to the need to compare vast amounts of new information against previously archived data, often failing to identify unchecked duplicate copies.
Innovation Solution
The method involves determining archive storage boundaries based on data mining and pattern recognition, selecting a preliminary boundary, performing sample duplication checks, and comparing results to optimize the storage boundary selection, thereby reducing duplicate storage and improving archival efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional methods compare each piece of newly received information to the vast amount of previously archived information, then duplicate copies can be identified, but processing time and resource consumption increase significantly
Solution Approach 1:
The patent divides the archival storage system into multiple segments or zones based on data characteristics, access patterns, and duplication probability. By segmenting the archive into logical boundaries, the system can perform duplication checks only within relevant segments rather than scanning the entire archive, significantly reducing processing time while maintaining duplicate identification accuracy.
Solution Approach 2:
The patent applies partial action by performing duplication checks on a subset of the archive rather than the complete archive. It uses sampling techniques and selective boundary determination to check only the necessary portions of archived data, achieving sufficient duplicate detection without the excessive resource consumption of exhaustive comparison.
2Reliability
If the archive storage boundary is set to cover the entire archived information, then all duplicate copies can be detected, but processing resources are excessively occupied
Solution Approach 1:
The patent implements dynamic archive storage boundaries that can be adjusted based on data characteristics, access patterns, and duplication probabilities. The boundaries are not fixed but adapt to changing conditions, allowing the system to optimize between detection completeness and processing efficiency by expanding or contracting the check scope as needed.
Solution Approach 2:
The patent changes key parameters such as archive storage boundary size, check depth, and sampling rate based on data analysis results. By dynamically adjusting these parameters, the system can achieve high duplicate detection rates when necessary while maintaining fast processing speeds during normal operations, thus resolving the contradiction between completeness and efficiency.
3Productivity
If the archive storage boundary is set to a small scope, then processing speed improves, but duplicate copies outside the boundary are not identified
Solution Approach 1:
The patent performs preliminary analysis of data characteristics, access patterns, and potential duplication sources before setting archive storage boundaries. This preliminary action allows the system to pre-identify high-probability duplicate regions and set boundaries that focus on these areas, ensuring both fast processing and accurate duplicate detection without needing to scan the entire archive.
Solution Approach 2:
The patent implements feedback mechanisms that continuously monitor archival operations, duplicate detection results, and data access patterns. Based on this feedback, the system dynamically adjusts archive storage boundaries to optimize the balance between processing speed and duplicate identification accuracy, learning from past performance to improve future boundary selections.
Data Source
AI summary
Archive systems and methods are presented. In one embodiment, an archival information storage configuration method comprises: performing an information accessing process including determining if the information is associated with an archive process; and performing an archive storage boundary determination process including establishing archive storage boundaries based upon characteristics indicating potential sharing of the information and potential impacts on performance of archival storage operations. In one exemplary implementation, the archive storage boundary determination process comprises: performing an information mining process including identifying an indication the information is potentially shared; and performing an archival boundary selection process including selecting an archive storage boundary based in at least part upon results of the information mining process.


