Non-disruptive Storage Controller Replacement via Cross-Cluster HA

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing storage cluster configurations face challenges in maintaining uninterrupted access to data during storage controller replacement within cross-cluster redundancy configurations, particularly in ensuring continuous operational continuity and minimizing downtime due to hardware or software failures.

Innovation Solution

The implementation of a cross-cluster redundancy configuration using High Availability (HA) pairs with synchronous data mirroring and NVRAM write cache replication, allowing for non-disruptive storage controller replacement by enabling one HA node to take over the operations of another, both within and across clusters, thereby maintaining data availability and minimizing downtime.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of repair

If storage controller replacement is performed in traditional storage cluster configurations, then hardware maintenance and upgrades can be conducted, but operational continuity is interrupted and data access is disrupted

Engineering Contradiction:
Improvestorage controller replacementVSAvoidoperational continuity
Core Design Contradiction:
Ease of repairVSReliability

Solution Approach 1:

The system performs preliminary actions by replicating data to a second storage controller before the replacement occurs. The second controller is pre-configured with redundant data and ready to immediately take over when the first controller is replaced, eliminating service interruption.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements local quality by creating a localized redundant copy of data specifically on the second storage controller within the same storage cluster. This local redundancy enables failover without requiring external systems, maintaining operational continuity during controller replacement.

Inventive Principle:
Principle #3Local quality

2Reliability

If cross-cluster data redundancy is implemented, then protection against site-wide failures is improved, but system complexity increases

Engineering Contradiction:
Improveprotection against site-wide failuresVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the redundancy function into two levels: intra-cluster redundancy through HA pairs for local failover, and inter-cluster redundancy through data replication for site-wide protection. This segmentation allows each layer to handle specific failure scenarios independently, managing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses storage arrays as intermediary components that facilitate data replication between storage controllers across clusters. These intermediaries manage the complexity of cross-cluster redundancy by providing standardized interfaces and automated replication processes.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11422908B2Non-disruptive controller replacement in a cross-cluster redundancy configuration
Publication Date: 2022.08.23 NETAPP INC
  • US11422908B2 patent drawing
  • US11422908B2 patent drawing
  • US11422908B2 patent drawing

AI summary

During a storage redundancy giveback from a first node to a second node following a storage redundancy takeover from the second node by the first node, the second node is initialized in part by receiving a node identification indicator from the second node. The node identification indicator is included in a node advertisement message sent by the second node during a giveback wait phase of the storage redundancy giveback. The node identification indicator includes an intra-cluster node connectivity identifier that is used by the first node to determine whether the second node is an intra-cluster takeover partner. In response to determining that the second node is an intra-cluster takeover partner, the first node completes the giveback of storage resources to the second node.