Remote Replication Bandwidth Prioritization for Storage SLO Compliance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional storage systems fail to ensure that remote replication meets Service Level Objectives (SLOs) during failover events, as they do not account for performance targets or relative business priorities of applications, leading to inefficient bandwidth usage and potential data loss.

Innovation Solution

A method that monitors network traffic characteristics and predicts changes in application demand and network state to dynamically manage replication, prioritizing critical applications and adjusting execution to meet performance targets and RTO/RPO objectives by using statistical models and business priority considerations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional storage systems perform remote replication without considering application priorities, then all applications receive equal bandwidth allocation, but critical applications cannot meet their Service Level Objectives during network congestion

Engineering Contradiction:
ImproveSLO complianceVSAvoidbandwidth utilization efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies local quality by differentiating bandwidth allocation based on application priorities. Critical applications receive higher bandwidth allocation while less critical applications receive lower allocation during network congestion. This is achieved through monitoring application performance metrics and dynamically adjusting replication bandwidth distribution to ensure SLO compliance for priority applications without completely starving lower-priority applications.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically changes the bandwidth allocation parameter based on real-time network conditions and application priorities. When network congestion is detected, the system adjusts the replication bandwidth parameter to favor critical applications. This parameter adjustment is continuous and adaptive, allowing the system to respond to changing conditions while maintaining overall system productivity.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the system prioritizes critical applications during replication, then SLOs are met for mission-critical apps, but less critical applications may experience delayed replication

Engineering Contradiction:
ImproveRTO/RPO achievementVSAvoidreplication delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by allocating replication bandwidth selectively rather than uniformly. During network congestion, the system provides excessive bandwidth to critical applications to ensure their RTO/RPO objectives are met, while accepting that less critical applications will receive partial bandwidth allocation resulting in delayed replication. This selective action ensures that the most important replication objectives are achieved without completely halting replication for lower-priority applications.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If the system monitors and predicts application demand dynamically, then bandwidth can be allocated optimally, but system complexity increases

Engineering Contradiction:
Improvebandwidth allocation efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms by continuously monitoring application performance metrics, network conditions, and replication status. This feedback information is used to dynamically adjust bandwidth allocation decisions. The system observes the outcomes of previous allocation decisions and adapts future allocations based on this feedback, enabling optimal bandwidth utilization without requiring overly complex predictive models. The feedback loop balances the need for dynamic optimization with acceptable system complexity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11128708B2Managing remote replication in storage systems
Publication Date: 2021.09.21 EMC IP HLDG CO LLC
  • US11128708B2 patent drawing
  • US11128708B2 patent drawing
  • US11128708B2 patent drawing

AI summary

A method is used in managing remote replication in storage systems. The method monitors network traffic characteristics of a network. The network enables communication between a first storage system and a second storage system. The method predicts a change in at least one of an application demand of an application of a set of applications executing on the first storage server and a network state of the network, where the set of applications have been identified for performing a replication to the second storage system. Based on the prediction, the method dynamically manages replication of the set of applications in accordance with a performance target associated with each application.