Secondary Backup Device Data Segmentation for De-duplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data backup solutions, particularly in the field of Purpose Built Backup Appliances (PBBA), face challenges with de-duplication efficiency due to high computational intensity and cost, where expensive high-end CPUs are required, and software-based solutions lack performance and wide applicability.
Innovation Solution
A method and device for data backup that involves a secondary backup device which segments target data, generates data fingerprints, and communicates these fingerprints to a primary backup device for de-duplication, reducing the computational load on the primary device by pre-processing new data segments and removing duplicates, thereby enhancing de-duplication efficiency and performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If de-duplication is performed using high-end CPUs, then de-duplication performance is improved, but device cost increases significantly
Solution Approach 1:
The patent divides the data backup system into two parts: a primary backup device and a secondary backup device. The secondary device performs computationally intensive de-duplication operations on data segments before they reach the primary device. This segmentation transfers the heavy computational workload from the expensive primary device to a more economical secondary device, resolving the contradiction between performance and cost.
Solution Approach 2:
The secondary backup device acts as an intermediary between the data source and the primary backup device. It pre-processes data by generating fingerprints and identifying duplicates before the data enters the primary system. This intermediary approach allows the primary device to focus only on actual new data, reducing its computational burden and allowing the use of more cost-effective hardware while maintaining high de-duplication performance.
2Ease of manufacture
If de-duplication is performed using software-based solutions, then device cost is reduced, but processing performance deteriorates
Solution Approach 1:
The patent creates a fingerprint copy of each data segment rather than working with the actual data. By generating a compact fingerprint representation and comparing these fingerprints to identify duplicates, the system achieves efficient de-duplication with minimal computational overhead. This copying approach maintains high performance while using cost-effective hardware, as the fingerprint operations are much lighter than full data processing.
3Productivity
If data segmentation and fingerprint generation are performed at the primary backup device, then de-duplication can be executed, but the primary device becomes a performance bottleneck
Solution Approach 1:
The secondary backup device performs preliminary actions by segmenting data and generating fingerprints before the data reaches the primary backup device. This pre-processing eliminates the bottleneck by completing the computationally intensive tasks in advance, allowing the primary device to simply receive and store the filtered data without performing complex operations itself.
Data Source
AI summary
Embodiments of the present disclosure provide a device for data backup comprising: a secondary backup device coupled to a primary backup device, the secondary backup device further comprising: data segmentation unit operable to divide target data to be backed up into a plurality of data segments; data fingerprint generation unit operable to generate a corresponding data fingerprint for each data segment from a plurality of data segments, and providing the data fingerprint to the primary backup device for backing up the target data at the primary backup device, wherein the data fingerprint is a mapped data segment of a length less than a corresponding data segment length.


