Aggregate Failover Data Structure for Cluster Storage Node Resilience
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current cluster storage systems face limitations in effectively managing node failures, as the workload of a failed primary node is often transferred to a single partner node, which can lead to reduced performance and inadequate protection against multiple node failures.
Innovation Solution
Implementing a takeover monitor module with an aggregate failover data structure (AFDS) that distributes the workload of a failed primary node across multiple partner nodes on a per aggregate basis, ensuring continued data servicing even after multiple node failures by specifying ordered lists of partner nodes to take over aggregate subsets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If the workload of a failed primary node is transferred to a single partner node, then the takeover process is simple, but the performance is reduced and protection against multiple node failures is inadequate
Solution Approach 1:
The patent segments the aggregate set of a failed primary node into multiple subset aggregates, each assigned to different partner nodes. This segmentation allows the workload to be distributed across multiple nodes rather than concentrated on a single node, improving reliability while maintaining manageable complexity through structured assignment rules.
Solution Approach 2:
The patent introduces a new dimension to the failover mechanism by implementing ordered lists of partner nodes for each subset aggregate. This creates a hierarchical failover structure where multiple levels of redundancy are established, transforming the single-dimension failover (one primary to one partner) into a multi-dimensional redundancy system.
2Device complexity
If the workload of a failed primary node is transferred to a single partner node, then the takeover process is simple, but the performance is reduced
Solution Approach 1:
By segmenting the workload into subset aggregates assigned to different partner nodes, the system distributes the processing burden across multiple nodes. This prevents any single node from becoming a performance bottleneck while maintaining organized and manageable takeover procedures.
Solution Approach 2:
The patent merges the capabilities of multiple partner nodes to service different subset aggregates of a failed primary node. This combining of resources across multiple nodes improves overall system performance by utilizing the collective capacity of the partner nodes rather than overloading a single node.
3Reliability
If multiple partner nodes take over subset aggregates of a failed primary node, then the resilience is enhanced, but the device complexity increases
Solution Approach 1:
The patent implements preliminary action by pre-establishing ordered lists of partner nodes for each subset aggregate before any failure occurs. This pre-configuration of failover relationships simplifies the actual takeover process, as the system only needs to follow the predetermined assignments rather than making complex decisions during failure events.
Solution Approach 2:
The system implements self-service through automated failover mechanisms that use the pre-configured aggregate failover data structure to automatically assign subset aggregates to appropriate partner nodes. This automation reduces the need for manual intervention and complex coordination, allowing the system to manage its own failover processes efficiently.
Data Source
AI summary
A cluster comprises a plurality of nodes that access a shared storage, each node having two or more partner nodes. A primary node may own a plurality of aggregate sub-sets in the shared storage. Upon failure of the primary node, each partner node may take over ownership of an aggregate sub-set according to an aggregate failover data structure (AFDS). The AFDS may specify, an ordered data structure of two or more partner nodes to take over each aggregate sub-set, the ordered data structure comprising at least a first-ordered partner node assigned to take over the aggregate sub-set upon failure of the primary node and a second-ordered partner node assigned to take over the aggregate sub-set upon failure of the primary node and the first-ordered partner node. The additional workload of the failed primary node is distributed among two or more partner nodes and protection for multiple node failures is provided.


