Virtual Storage Controller Failover in Scale-Out Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional high availability storage systems face unavailability when both primary and secondary nodes fail, and manual intervention is required for recovery, leading to downtime and limited scalability.

Innovation Solution

Implementing a dynamic failover process that designates a new secondary access node in response to node failures, allowing automatic handling of failures without manual intervention, and enabling failover without disrupting data access or requiring data relocation, using a system with virtual access nodes and scalable redundancy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a conventional storage system uses only primary and secondary nodes for high availability, then the system can provide basic failover capability, but the system becomes unavailable when both primary and secondary nodes fail, and manual intervention is required for recovery

Engineering Contradiction:
Improvesystem availabilityVSAvoidfailure tolerance capacity
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the access node role into multiple independent nodes (first access node, second access node, third access node) rather than relying on a fixed primary-secondary pair. This segmentation allows the system to tolerate multiple failures by distributing the access function across several nodes, each capable of independently servicing requests.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic role assignment where access node designations are not fixed but can be changed based on operational needs and failure conditions. The system can dynamically reassign access node roles to maintain availability, transitioning from a static primary-secondary model to a flexible multi-node dynamic model.

Inventive Principle:
Principle #15Dynamics

2Reliability

If manual intervention is required for node failure recovery, then the system structure can be simpler, but the mean time to repair (MTTR) increases and system availability decreases

Engineering Contradiction:
Improvesystem availabilityVSAvoidmean time to repair
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements self-service failover capability where the storage system automatically detects node failures and performs failover operations without requiring manual administrator intervention. The system monitors node health, detects failures, and automatically reconfigures access node designations to maintain continuity, thereby reducing MTTR and improving availability.

Inventive Principle:
Principle #25Self-service

3Reliability

If failover requires data relocation to maintain availability, then data integrity can be preserved, but data access is disrupted during the relocation process

Engineering Contradiction:
Improvedata integrityVSAvoiddata access continuity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements preliminary action by pre-designating multiple access nodes (first, second, and third access nodes) before any failure occurs. This advance preparation ensures that when a failure happens, the system can immediately switch to the next available access node without needing to relocate data or disrupt ongoing access operations, thereby maintaining both integrity and continuity.

Inventive Principle:
Principle #10Preliminary action

4Device complexity

If the storage system uses fixed primary and secondary access nodes, then the system structure is simpler, but the system cannot provide high availability when multiple nodes fail simultaneously

Engineering Contradiction:
Improveaccess node configurationVSAvoidmulti-node failure tolerance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the access function across multiple independent nodes (first access node, second access node, third access node) rather than using a fixed primary-secondary pair. This segmentation increases the system's ability to tolerate multiple simultaneous failures while maintaining a relatively simple overall structure, as each node operates independently with similar functionality.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240311246A1High availability using virtual storage controllers in a scale out storage cluster
Publication Date: 2024.09.19 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240311246A1 patent drawing
  • US20240311246A1 patent drawing
  • US20240311246A1 patent drawing

AI summary

Examples are provided for a method of providing access to data of a data center. In one aspect, the method comprises storing a unit of data to each of a plurality of data nodes of a data center, designating a first node of the data center as a primary access node for the unit of data, the primary access node being configured to service access requests to the unit of data using one or more of the plurality of data nodes, determining that the first node is not available, and performing a failover process by reconfiguring a second node of the data center as the primary access node for the unit of data.