Distributed Grid Server Deduplication for Backup Performance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional deduplication systems face availability issues due to reliance on a single compute server for site-wide deduplication, leading to increased backup times and storage capacity consumption, and are inefficient in processing large backup images, especially when the deduplication appliance fails during the backup process.

Innovation Solution

A distributed grid server network architecture that allows parallel deduplication across multiple servers, using zone stamps to identify and deduplicate data segments, and enables failover models to maintain availability and reduce storage and bandwidth requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single compute server is used for site-wide deduplication, then the system architecture is simple, but the backup time increases and storage capacity consumption increases

Engineering Contradiction:
Improvesystem architectureVSAvoidbackup time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent divides the data stream into multiple segments and distributes them across multiple grid servers for parallel deduplication processing. Each server handles a portion of the data independently, enabling concurrent processing that reduces overall backup time while maintaining system scalability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a single-server sequential processing model to a multi-server parallel processing architecture. By adding the dimension of multiple processing nodes working simultaneously, the system achieves faster backup times without proportionally increasing complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If a single compute server is used for site-wide deduplication, then the system architecture is simple, but storage capacity consumption increases

Engineering Contradiction:
Improvesystem architectureVSAvoidstorage capacity consumption
Core Design Contradiction:
Device complexityVSQuantity of substance

Solution Approach 1:

The patent segments the data processing workload across multiple grid servers, allowing distributed deduplication operations. This enables more efficient identification and elimination of duplicate data blocks across the entire dataset, reducing overall storage capacity consumption compared to single-server processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a controlled environment where deduplication operations are performed across multiple isolated processing units (grid servers), each maintaining its own deduplication context. This prevents redundant processing and ensures comprehensive deduplication coverage, minimizing storage capacity consumption.

Inventive Principle:
Principle #39Inert atmosphere (Inert environment)

3Quantity of substance

If multiple disk storage units are added to handle increasing backup data, then storage capacity increases, but backup time exceeds service level agreement limits

Engineering Contradiction:
Improvestorage capacityVSAvoidbackup time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent divides the incoming data stream into multiple segments that can be processed in parallel by multiple grid servers. This segmentation allows the system to utilize multiple storage units simultaneously for deduplication operations, maintaining backup times within service level agreement limits even as storage capacity scales.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary segmentation and distribution of data to multiple grid servers before the actual deduplication process. This preparation enables concurrent processing of multiple data portions, ensuring that backup operations complete within required timeframes regardless of the total data volume.

Inventive Principle:
Principle #10Preliminary action

4Device complexity

If a single deduplication appliance is used, then the system is simple to manage, but the system fails completely when the appliance fails

Engineering Contradiction:
Improvesystem managementVSAvoidsystem availability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent distributes deduplication functionality across multiple independent grid servers rather than relying on a single appliance. This segmentation creates redundancy where individual server failures do not cause complete system failure, improving reliability while maintaining manageable system operations through standardized server interfaces.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a distributed architecture where multiple grid servers are prepared in advance to handle deduplication workloads. If one server fails, others are already positioned and configured to take over its workload, providing fault tolerance and maintaining system availability without complex failover mechanisms.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

5Device complexity

If a single front-end compute node is used, then the system architecture is simple, but parallel processing capability is limited

Engineering Contradiction:
Improvesystem architectureVSAvoiddeduplication processing speed
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent segments the deduplication processing workload into multiple independent tasks that can be executed in parallel across multiple grid servers. Each server processes assigned data segments concurrently, dramatically increasing overall deduplication throughput compared to a single front-end compute node while maintaining architectural simplicity through standardized processing units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a single-node sequential processing architecture to a multi-node parallel processing architecture. By introducing multiple processing dimensions (multiple grid servers operating simultaneously), the system achieves enhanced productivity and deduplication processing speed without proportionally increasing management complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11182345B2Parallelizing and deduplicating backup data
Publication Date: 2021.11.23 EXAGRID SYST
  • US11182345B2 patent drawing
  • US11182345B2 patent drawing
  • US11182345B2 patent drawing

AI summary

A method, a system, and a computer program product for performing a backup of data are disclosed. A grid server in a plurality of grid servers is selected for deduplicating a segment of data in a plurality of segments of data contained within a data stream. The segment of data is forwarded to the selected grid server for deduplication. A zone contained within the forwarded segment of data is deduplicated using the selected server. The deduplication is performed based on a listing of a plurality of zone stamps. Each zone stamp in the plurality of zone stamps represents a zone in a plurality of zones deduplicated by at least one server in the plurality of grid servers.