Virtual Storage Controller Failover in Scale-Out Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional high availability storage systems face unavailability when both primary and secondary nodes fail, and manual intervention is required for recovery, leading to downtime and limited scalability.
Innovation Solution
Implementing a dynamic failover process that designates a new secondary access node in response to node failures, allowing automatic handling of failures without manual intervention, and enabling failover without disrupting data access or requiring data relocation, using a system with virtual access nodes and scalable redundancy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a conventional storage system uses only primary and secondary nodes for high availability, then the system can provide basic failover capability, but the system becomes unavailable when both primary and secondary nodes fail, and manual intervention is required for recovery
Solution Approach 1:
The patent segments the access node role into multiple independent nodes (first access node, second access node, third access node) rather than relying on a fixed primary-secondary pair. This segmentation allows the system to tolerate multiple failures by distributing the access function across several nodes, each capable of independently servicing requests.
Solution Approach 2:
The patent implements dynamic role assignment where access node designations are not fixed but can be changed based on operational needs and failure conditions. The system can dynamically reassign access node roles to maintain availability, transitioning from a static primary-secondary model to a flexible multi-node dynamic model.
2Reliability
If manual intervention is required for node failure recovery, then the system structure can be simpler, but the mean time to repair (MTTR) increases and system availability decreases
Solution Approach 1:
The patent implements self-service failover capability where the storage system automatically detects node failures and performs failover operations without requiring manual administrator intervention. The system monitors node health, detects failures, and automatically reconfigures access node designations to maintain continuity, thereby reducing MTTR and improving availability.
3Reliability
If failover requires data relocation to maintain availability, then data integrity can be preserved, but data access is disrupted during the relocation process
Solution Approach 1:
The patent implements preliminary action by pre-designating multiple access nodes (first, second, and third access nodes) before any failure occurs. This advance preparation ensures that when a failure happens, the system can immediately switch to the next available access node without needing to relocate data or disrupt ongoing access operations, thereby maintaining both integrity and continuity.
4Device complexity
If the storage system uses fixed primary and secondary access nodes, then the system structure is simpler, but the system cannot provide high availability when multiple nodes fail simultaneously
Solution Approach 1:
The patent segments the access function across multiple independent nodes (first access node, second access node, third access node) rather than using a fixed primary-secondary pair. This segmentation increases the system's ability to tolerate multiple simultaneous failures while maintaining a relatively simple overall structure, as each node operates independently with similar functionality.
Data Source
AI summary
Examples are provided for a method of providing access to data of a data center. In one aspect, the method comprises storing a unit of data to each of a plurality of data nodes of a data center, designating a first node of the data center as a primary access node for the unit of data, the primary access node being configured to service access requests to the unit of data using one or more of the plurality of data nodes, determining that the first node is not available, and performing a failover process by reconfiguring a second node of the data center as the primary access node for the unit of data.


