Distributed Grid Server Deduplication for Backup Performance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional deduplication systems face availability issues due to reliance on a single compute server for site-wide deduplication, leading to increased backup times and storage capacity consumption, and are inefficient in processing large backup images, especially when the deduplication appliance fails during the backup process.
Innovation Solution
A distributed grid server network architecture that allows parallel deduplication across multiple servers, using zone stamps to identify and deduplicate data segments, and enables failover models to maintain availability and reduce storage and bandwidth requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single compute server is used for site-wide deduplication, then the system architecture is simple, but the backup time increases and storage capacity consumption increases
Solution Approach 1:
The patent divides the data stream into multiple segments and distributes them across multiple grid servers for parallel deduplication processing. Each server handles a portion of the data independently, enabling concurrent processing that reduces overall backup time while maintaining system scalability.
Solution Approach 2:
The patent transitions from a single-server sequential processing model to a multi-server parallel processing architecture. By adding the dimension of multiple processing nodes working simultaneously, the system achieves faster backup times without proportionally increasing complexity.
2Device complexity
If a single compute server is used for site-wide deduplication, then the system architecture is simple, but storage capacity consumption increases
Solution Approach 1:
The patent segments the data processing workload across multiple grid servers, allowing distributed deduplication operations. This enables more efficient identification and elimination of duplicate data blocks across the entire dataset, reducing overall storage capacity consumption compared to single-server processing.
Solution Approach 2:
The patent creates a controlled environment where deduplication operations are performed across multiple isolated processing units (grid servers), each maintaining its own deduplication context. This prevents redundant processing and ensures comprehensive deduplication coverage, minimizing storage capacity consumption.
3Quantity of substance
If multiple disk storage units are added to handle increasing backup data, then storage capacity increases, but backup time exceeds service level agreement limits
Solution Approach 1:
The patent divides the incoming data stream into multiple segments that can be processed in parallel by multiple grid servers. This segmentation allows the system to utilize multiple storage units simultaneously for deduplication operations, maintaining backup times within service level agreement limits even as storage capacity scales.
Solution Approach 2:
The patent performs preliminary segmentation and distribution of data to multiple grid servers before the actual deduplication process. This preparation enables concurrent processing of multiple data portions, ensuring that backup operations complete within required timeframes regardless of the total data volume.
4Device complexity
If a single deduplication appliance is used, then the system is simple to manage, but the system fails completely when the appliance fails
Solution Approach 1:
The patent distributes deduplication functionality across multiple independent grid servers rather than relying on a single appliance. This segmentation creates redundancy where individual server failures do not cause complete system failure, improving reliability while maintaining manageable system operations through standardized server interfaces.
Solution Approach 2:
The patent implements a distributed architecture where multiple grid servers are prepared in advance to handle deduplication workloads. If one server fails, others are already positioned and configured to take over its workload, providing fault tolerance and maintaining system availability without complex failover mechanisms.
5Device complexity
If a single front-end compute node is used, then the system architecture is simple, but parallel processing capability is limited
Solution Approach 1:
The patent segments the deduplication processing workload into multiple independent tasks that can be executed in parallel across multiple grid servers. Each server processes assigned data segments concurrently, dramatically increasing overall deduplication throughput compared to a single front-end compute node while maintaining architectural simplicity through standardized processing units.
Solution Approach 2:
The patent transitions from a single-node sequential processing architecture to a multi-node parallel processing architecture. By introducing multiple processing dimensions (multiple grid servers operating simultaneously), the system achieves enhanced productivity and deduplication processing speed without proportionally increasing management complexity.
Data Source
AI summary
A method, a system, and a computer program product for performing a backup of data are disclosed. A grid server in a plurality of grid servers is selected for deduplicating a segment of data in a plurality of segments of data contained within a data stream. The segment of data is forwarded to the selected grid server for deduplication. A zone contained within the forwarded segment of data is deduplicated using the selected server. The deduplication is performed based on a listing of a plurality of zone stamps. Each zone stamp in the plurality of zone stamps represents a zone in a plurality of zones deduplicated by at least one server in the plurality of grid servers.


