Workload Prediction for Data Restore Site Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems lack the ability to intelligently select the best data protection site for restoring data, especially when the preferred disaster recovery site experiences low performance, impacting normal backup operations and service level agreements, and fail to optimize restore operations for mobile users who may not have a configured disaster recovery site or have it geographically closest.

Innovation Solution

An analytic engine collects workload history data to predict IO and CPU loads, sharing these predictions across data protection sites to redirect restore requests to the least busy site, using machine learning techniques to determine future workload status and balance load across replication sites.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If restore requests are directed to the preferred disaster recovery site, then data availability is improved, but normal backup operations are impacted due to low performance

Engineering Contradiction:
Improvedata availabilityVSAvoidbackup operation performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent introduces a workload prediction mechanism as an intermediary between restore requests and DR sites. The system predicts future workload at each DR site and uses these predictions to intelligently route restore requests, thereby mediating between the need for data availability and the need to protect backup operations from performance degradation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary workload prediction before directing restore requests. By analyzing historical workload data and predicting future workload status of DR sites, the system proactively identifies optimal restore destinations that will minimize impact on backup operations, rather than reactively responding to performance issues after they occur

Inventive Principle:
Principle #10Preliminary action

2Reliability

If multiple replicas are maintained across different data protection sites, then data availability is improved, but the complexity of selecting the optimal restore site increases

Engineering Contradiction:
Improvedata availabilityVSAvoidrestore site selection complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent changes the parameter used for site selection from static criteria (such as geographic proximity or configured preference) to dynamic workload predictions. By continuously updating workload predictions based on historical data and using these predictions to guide restore requests, the system simplifies the selection process while maintaining optimality across multiple replicas

Inventive Principle:
Principle #35Parameter changes

3Speed

If restore operations are performed at the geographically closest DR site, then response time is improved, but service level agreements may not be met due to high workload

Engineering Contradiction:
Improverestore response timeVSAvoidservice level agreement compliance
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The system transitions from static geographic-based site selection to dynamic workload-based selection. By continuously predicting workload and adapting restore request routing accordingly, the system can dynamically choose between geographically closer but busy sites and farther but idle sites, ensuring both fast response times and SLA compliance

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11693743B2Method to optimize restore based on data protection workload prediction
Publication Date: 2023.07.04 EMC IP HLDG CO LLC
  • US11693743B2 patent drawing
  • US11693743B2 patent drawing
  • US11693743B2 patent drawing

AI summary

An intelligent method of selecting a data recovery site upon receiving a data recovery request. The backup system collects historical activity data of the storage system to identify work load of every data recovery site. A predicted activity load for each data recovery site is then generated using the collected data. When a request for data recovery is received, the system first identifies which data recovery site has copies of the files to be recovered. Then it uses the predicted work load for these data recovery sites to determine whether to use a geographically local site or a site that may be remote geographically, but has a lower work load.