Multi-site Storage Snapshot Retention and Failover Reconfiguration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multi-site distributed storage systems face challenges in maintaining common snapshots between secondary and tertiary storage sites during a primary storage site failure, leading to time-consuming and bandwidth-intensive baseline data transfers for resumption of protection configurations.

Innovation Solution

Implementing an asynchronous mirroring policy to ensure common snapshots are maintained between secondary and tertiary storage sites, enabling automatic unplanned failover and realignment of protection configurations without baseline data transfers, and automatically reconfiguring replication relationships post-failure.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If baseline data transfer is performed between secondary and tertiary storage sites after primary storage failure, then data consistency is restored, but network bandwidth consumption increases and recovery time extends

Engineering Contradiction:
Improvedata consistencyVSAvoidnetwork bandwidth consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system performs preliminary actions by maintaining asynchronous replication relationships and snapshot copies between secondary and tertiary storage sites before the primary storage failure occurs. This preliminary setup ensures that when failover happens, data can be restored without requiring extensive baseline transfers, thus reducing network bandwidth consumption during recovery.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent utilizes snapshot copies as lightweight replicas of the original data. Instead of transferring complete baseline data between storage sites, the system creates and maintains snapshot copies that can be quickly activated during failover, significantly reducing the network bandwidth required for data consistency restoration.

Inventive Principle:
Principle #26Copying

2Loss of time

If manual user intervention is used to restore operations after storage site failure, then system complexity is reduced, but recovery time increases

Engineering Contradiction:
Improverecovery timeVSAvoidautomatic failover capability
Core Design Contradiction:
Loss of timeVSExtent of automation

Solution Approach 1:

The system implements self-service capabilities through automatic failover mechanisms. When the primary storage site fails, the secondary storage site automatically takes over without requiring manual user intervention. The system also automatically initiates realignment and reconfiguration of protection configurations, enabling rapid recovery while maintaining high automation levels.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system employs feedback mechanisms to monitor the health status of storage sites and automatically trigger failover procedures when failures are detected. This closed-loop control enables the system to respond to failures autonomously, reducing recovery time without sacrificing automation.

Inventive Principle:
Principle #23Feedback

3Reliability

If synchronous replication is used from primary to secondary storage site, then data availability is improved, but network bandwidth consumption increases

Engineering Contradiction:
Improvedata availabilityVSAvoidnetwork bandwidth consumption
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system applies different replication strategies to different data paths based on local requirements. Synchronous replication is used from primary to secondary storage site to ensure data availability, while asynchronous replication is used from primary to tertiary storage site to reduce network bandwidth consumption. This localized optimization allows the system to achieve reliability where needed while minimizing bandwidth usage elsewhere.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically changes replication parameters based on operational conditions. The replication mode (synchronous or asynchronous) and update schedules are adjusted according to the specific data path and requirements, allowing optimization of both data availability and network bandwidth consumption for different replication relationships.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If protection configuration realignment is performed automatically after failover, then operational continuity is maintained, but system complexity increases

Engineering Contradiction:
Improveoperational continuityVSAvoidsystem configuration complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements a universal realignment mechanism that automatically handles multiple configuration tasks after failover. The same automated process manages both the activation of protection configurations and the reconfiguration of replication relationships, reducing the need for separate manual intervention steps and maintaining operational continuity despite the inherent complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250110837A1Methods and multi-site systems to provide recovery point objective (RPO) protection, snapshot retention between secondary storage site and tertiary storage site, and automatically initiating realignment and reconfiguration of a protection configuration from the secondary storage site to the tertiary storage site upon primary storage site failure
Publication Date: 2025.04.03 NETAPP INC
  • US20250110837A1 patent drawing
  • US20250110837A1 patent drawing
  • US20250110837A1 patent drawing

AI summary

Multi-site distributed storage systems and computer-implemented methods are described for providing common snapshot retention and automatic fanout reconfiguration for an asynchronous leg after a failure event that causes a failover from a primary storage site to a secondary storage site. A computer-implemented method comprises providing an asynchronous replication relationship with an asynchronous update schedule from one or more storage objects of the first storage node to one or more replicated storage objects of the third storage node, creating a snapshot copy of the one or more storage objects of the first storage node, transferring the snapshot copy to the third storage node based on an asynchronous mirror policy, and intercepting the snapshot create operation on the primary storage site and transferring the snapshot copy to the second storage node to provide a common snapshot between the second storage node and the third storage node.