Non-disruptive Storage Controller Replacement via Cross-Cluster HA
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing storage cluster configurations face challenges in maintaining uninterrupted access to data during storage controller replacement within cross-cluster redundancy configurations, particularly in ensuring continuous operational continuity and minimizing downtime due to hardware or software failures.
Innovation Solution
The implementation of a cross-cluster redundancy configuration using High Availability (HA) pairs with synchronous data mirroring and NVRAM write cache replication, allowing for non-disruptive storage controller replacement by enabling one HA node to take over the operations of another, both within and across clusters, thereby maintaining data availability and minimizing downtime.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of repair
If storage controller replacement is performed in traditional storage cluster configurations, then hardware maintenance and upgrades can be conducted, but operational continuity is interrupted and data access is disrupted
Solution Approach 1:
The system performs preliminary actions by replicating data to a second storage controller before the replacement occurs. The second controller is pre-configured with redundant data and ready to immediately take over when the first controller is replaced, eliminating service interruption.
Solution Approach 2:
The patent implements local quality by creating a localized redundant copy of data specifically on the second storage controller within the same storage cluster. This local redundancy enables failover without requiring external systems, maintaining operational continuity during controller replacement.
2Reliability
If cross-cluster data redundancy is implemented, then protection against site-wide failures is improved, but system complexity increases
Solution Approach 1:
The system segments the redundancy function into two levels: intra-cluster redundancy through HA pairs for local failover, and inter-cluster redundancy through data replication for site-wide protection. This segmentation allows each layer to handle specific failure scenarios independently, managing overall system complexity.
Solution Approach 2:
The patent uses storage arrays as intermediary components that facilitate data replication between storage controllers across clusters. These intermediaries manage the complexity of cross-cluster redundancy by providing standardized interfaces and automated replication processes.
Data Source
AI summary
During a storage redundancy giveback from a first node to a second node following a storage redundancy takeover from the second node by the first node, the second node is initialized in part by receiving a node identification indicator from the second node. The node identification indicator is included in a node advertisement message sent by the second node during a giveback wait phase of the storage redundancy giveback. The node identification indicator includes an intra-cluster node connectivity identifier that is used by the first node to determine whether the second node is an intra-cluster takeover partner. In response to determining that the second node is an intra-cluster takeover partner, the first node completes the giveback of storage resources to the second node.


