Intelligent Parallel Cloud Backup Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data recovery processes from cloud backups are time-consuming due to the limitations of Wide Area Network (WAN) speeds, resulting in significant downtime, especially when restoring large datasets like 10 Terabytes over a 1 Gb/s link which can take approximately 27 hours.

Innovation Solution

Implementing a data protection system that performs intelligent parallel recovery from multiple backup copies stored across different locations, optimizing data download by splitting data into segments and distributing them across various backup sites based on their download throughput, allowing simultaneous download from multiple locations to minimize overall recovery time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is recovered from a single cloud backup copy over WAN, then data protection is ensured, but recovery time becomes excessively long

Engineering Contradiction:
Improvedata protectionVSAvoidrecovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the backup copy into multiple data portions and distributes them across multiple cloud storage locations. During recovery, these segments are simultaneously downloaded from multiple sources in parallel, dramatically reducing total recovery time while maintaining data integrity through proper segment assembly

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent combines multiple cloud storage locations into a unified recovery system, where data portions from different locations work together to reconstruct the complete dataset. This merging of resources transforms a single-point bottleneck into a multi-path recovery architecture

Inventive Principle:
Principle #5Merging (Combining)

2Loss of time

If data is segmented and downloaded from multiple cloud locations in parallel, then recovery time is reduced, but system complexity increases

Engineering Contradiction:
Improverecovery timeVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent introduces a recovery manager as an intermediary component that coordinates the complex multi-location recovery process. This mediator handles segment tracking, download coordination, and assembly orchestration, shielding users from the underlying complexity while enabling parallel recovery operations

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from a single-dimension recovery approach (one backup location) to a multi-dimensional architecture by distributing data portions across multiple cloud locations. This dimensional expansion enables parallel operations without proportionally increasing user-facing complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If multiple backup copies are stored across different cloud locations, then recovery speed is improved, but storage costs increase

Engineering Contradiction:
Improverecovery speedVSAvoidstorage capacity
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

Instead of creating complete duplicate backups at multiple locations, the patent segments a single backup into portions distributed across multiple cloud storage locations. This segmentation strategy achieves multi-location recovery capability while using significantly less total storage capacity than full duplicates

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3974987B1Intelligent recovery from multiple cloud copies
Publication Date: 2023.06.28 EMC IP HLDG CO LLC
  • EP3974987B1 patent drawingFigure 1
  • EP3974987B1 patent drawingFigure 2
  • EP3974987B1 patent drawingFigure 3

AI summary

Data protection operations including recovery operations are disclosed. A recovery operation is performed by downloading the backup to be recovered from multiple identical copies of the backup. Based on factors such as throughput, an optimal amount of data can be downloaded from each of the multiple copies. The portions downloaded from the multiple copies are rebuild or reassembled once downloaded. The assembled backup is then presented to a target for recovery.