Storage Controller Automatic Switchover via RDMA Heartbeat

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing storage networks with single controller configurations lack local high availability, leading to slow and imprecise cross-cluster switchover operations in case of storage controller failures, resulting in disruptive client access to data.

Innovation Solution

Implementing automatic switchover techniques using remote direct memory access (RDMA) read operations and service processor traps to quickly detect storage controller failures and initiate switchover operations, ensuring non-disruptive client access by a surviving storage controller.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If single controller configurations are used to reduce cost, then device complexity is reduced, but reliability deteriorates due to lack of local high availability

Engineering Contradiction:
Improvecontroller configurationVSAvoidlocal high availability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces a remote storage controller as an intermediary that can detect failures and initiate switchover operations across clusters. This mediator enables automatic failover without requiring local high-availability controllers, thus maintaining reliability while using simple single-controller configurations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by having the remote storage controller continuously monitor the primary controller's health through heartbeat signals and service processor traps. When a failure is detected, the switchover operation is automatically initiated without waiting for manual intervention, ensuring continuous data access.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If manual switchover operations are used to ensure data access, then device complexity is reduced, but speed of failover deteriorates causing client access disruption

Engineering Contradiction:
Improveswitchover mechanismVSAvoidfailover speed
Core Design Contradiction:
Device complexityVSSpeed

Solution Approach 1:

The patent implements feedback mechanisms through heartbeat signals and service processor traps that continuously monitor the primary storage controller's status. When a failure is detected, this feedback triggers an automatic switchover operation, eliminating the need for manual intervention and reducing client access disruption.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The remote storage controller performs self-service by automatically detecting failures through monitored heartbeats and initiating switchover operations without external intervention. This automation speeds up failover while keeping the mechanism relatively simple.

Inventive Principle:
Principle #25Self-service

3Reliability

If cross-cluster remote detection techniques are used to detect storage controller failure, then reliability is improved, but speed of detection deteriorates due to use of timeouts and manual techniques

Engineering Contradiction:
Improvefailure detection reliabilityVSAvoidfailure detection speed
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The patent uses feedback mechanisms where the primary storage controller sends periodic heartbeat signals to the remote controller, and the remote controller monitors service processor traps. This continuous feedback loop enables rapid and reliable automatic failure detection across clusters, replacing slow timeout-based methods.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces mechanical/manual failure detection techniques with electronic monitoring mechanisms. Service processor traps and heartbeat signals provide automated, rapid detection of storage controller failures, eliminating the delays associated with manual detection methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11232004B2Implementing automatic switchover
Publication Date: 2022.01.25 NETAPP INC
  • US11232004B2 patent drawing
  • US11232004B2 patent drawing
  • US11232004B2 patent drawing

AI summary

One or more techniques and/or computing devices are provided for automatic switchover implementation. For example, a first storage controller, of a first storage cluster, may have a disaster recovery relationship with a second storage controller of a second storage cluster. In the event the first storage controller fails, the second storage controller may automatically switchover operation from the first storage controller to the second storage controller for providing clients with failover access to data previously accessible to the clients through the first storage controller. The second storage controller may detect, cross-cluster, a failure of the first storage controller utilizing remote direct memory access (RDMA) read operations to access heartbeat information, heartbeat information stored within a disk mailbox, and/or service processor traps. In this way, the second storage controller may efficiently detect failure of the first storage controller to trigger automatic switchover for non-disruptive client access to data.