Archive Storage Boundary Segmentation for Duplicate Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for identifying and reducing duplicate copies in archival storage systems are inefficient, leading to long delays and excessive resource occupation due to the need to compare vast amounts of new information against previously archived data, often failing to identify unchecked duplicate copies.

Innovation Solution

The method involves determining archive storage boundaries based on data mining and pattern recognition, selecting a preliminary boundary, performing sample duplication checks, and comparing results to optimize the storage boundary selection, thereby reducing duplicate storage and improving archival efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional methods compare each piece of newly received information to the vast amount of previously archived information, then duplicate copies can be identified, but processing time and resource consumption increase significantly

Engineering Contradiction:
Improveduplicate identification accuracyVSAvoidarchival processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent divides the archival storage system into multiple segments or zones based on data characteristics, access patterns, and duplication probability. By segmenting the archive into logical boundaries, the system can perform duplication checks only within relevant segments rather than scanning the entire archive, significantly reducing processing time while maintaining duplicate identification accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by performing duplication checks on a subset of the archive rather than the complete archive. It uses sampling techniques and selective boundary determination to check only the necessary portions of archived data, achieving sufficient duplicate detection without the excessive resource consumption of exhaustive comparison.

Inventive Principle:
Principle #16Partial or excessive action

2Reliability

If the archive storage boundary is set to cover the entire archived information, then all duplicate copies can be detected, but processing resources are excessively occupied

Engineering Contradiction:
Improveduplicate detection completenessVSAvoidarchival operation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements dynamic archive storage boundaries that can be adjusted based on data characteristics, access patterns, and duplication probabilities. The boundaries are not fixed but adapt to changing conditions, allowing the system to optimize between detection completeness and processing efficiency by expanding or contracting the check scope as needed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes key parameters such as archive storage boundary size, check depth, and sampling rate based on data analysis results. By dynamically adjusting these parameters, the system can achieve high duplicate detection rates when necessary while maintaining fast processing speeds during normal operations, thus resolving the contradiction between completeness and efficiency.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If the archive storage boundary is set to a small scope, then processing speed improves, but duplicate copies outside the boundary are not identified

Engineering Contradiction:
Improvearchival processing speedVSAvoidduplicate identification accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary analysis of data characteristics, access patterns, and potential duplication sources before setting archive storage boundaries. This preliminary action allows the system to pre-identify high-probability duplicate regions and set boundaries that focus on these areas, ensuring both fast processing and accurate duplicate detection without needing to scan the entire archive.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms that continuously monitor archival operations, duplicate detection results, and data access patterns. Based on this feedback, the system dynamically adjusts archive storage boundaries to optimize the balance between processing speed and duplicate identification accuracy, learning from past performance to improve future boundary selections.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9690789B2Archive systems and methods
Publication Date: 2017.06.27 COHESITY INC
  • US9690789B2 patent drawing
  • US9690789B2 patent drawing
  • US9690789B2 patent drawing

AI summary

Archive systems and methods are presented. In one embodiment, an archival information storage configuration method comprises: performing an information accessing process including determining if the information is associated with an archive process; and performing an archive storage boundary determination process including establishing archive storage boundaries based upon characteristics indicating potential sharing of the information and potential impacts on performance of archival storage operations. In one exemplary implementation, the archive storage boundary determination process comprises: performing an information mining process including identifying an indication the information is potentially shared; and performing an archival boundary selection process including selecting an archive storage boundary based in at least part upon results of the information mining process.