Multi-pathing Driver I/O Redirection for Storage Failover

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing failover protocols in enterprise storage arrays often result in costly and unnecessary downtime when connectivity is compromised, as they trigger cluster-wide failovers even for isolated node failures, leading to paused I/O processing and performance delays.

Innovation Solution

Implementing a method to redirect I/Os from a local host to an available remote host capable of delivering I/Os to the storage system, thereby avoiding the need for a cluster-wide failover protocol and maintaining access during connectivity failures or maintenance operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a failover protocol is implemented to ensure continuous access to storage arrays, then reliability is improved, but productivity deteriorates due to costly downtime and paused I/O processing

Engineering Contradiction:
Improvecontinuous access to storage arrayVSAvoidI/O processing throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the failover detection and response mechanism by node. Instead of a centralized failover protocol that affects the entire cluster, each node independently monitors its own connectivity to the storage array and executes failover only when necessary. This segmentation prevents cluster-wide I/O pauses when a single node experiences connectivity issues, thereby maintaining productivity while ensuring reliability for affected nodes.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by allowing different nodes in the cluster to have different failover states simultaneously. When one node experiences connectivity loss, only that specific node initiates failover to alternative paths, while other nodes continue normal I/O operations. This localized response maintains overall cluster productivity while ensuring continuous access for the affected node.

Inventive Principle:
Principle #3Local quality

2Reliability

If a cluster-wide failover protocol is triggered to handle connectivity failures, then reliability is improved, but loss of time increases due to paused I/O processing across all nodes

Engineering Contradiction:
Improveaccess to storage array during failureVSAvoidI/O pause duration during failover
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent divides the failover execution into node-specific operations rather than a synchronized cluster-wide event. Each node independently determines when to pause I/O and when to switch paths, minimizing the time that any given node's I/O is paused. Other nodes experience no interruption, significantly reducing overall loss of time while maintaining reliability for the failed node.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action by pre-configuring multiple active paths to the storage array before failures occur. When connectivity is lost on one node, the failover to pre-configured alternative paths can occur rapidly without requiring extensive reconfiguration or coordination with other nodes, thereby reducing the time I/O processing is paused.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If maintenance operations are performed on a primary host controller, then adaptability is improved, but productivity deteriorates due to cluster-wide access prevention

Engineering Contradiction:
Improvemaintenance capabilityVSAvoidstorage array access during maintenance
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments the maintenance operation impact to the specific node undergoing maintenance. When a primary host controller on one node is taken down for maintenance, only that node experiences connectivity loss and initiates failover. Other nodes in the cluster continue to access the storage array normally through their own primary controllers, maintaining productivity while allowing necessary maintenance to proceed on the affected node.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by allowing different nodes to operate in different states simultaneously. The node undergoing maintenance experiences failover to alternative paths, while other nodes maintain normal operation. This localized impact enables maintenance operations to proceed without sacrificing cluster-wide productivity.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9015519B2Method and system for cluster wide adaptive I/O scheduling by a multipathing driver
Publication Date: 2015.04.21 ARCTERA US LLC
  • US9015519B2 patent drawing
  • US9015519B2 patent drawing
  • US9015519B2 patent drawing

AI summary

A method and system for load balancing. The method includes determining that connectivity between a first host and a primary array controller of a storage system has failed. The first host is configured to send input/output messages (I/Os) to a storage system through a storage network fabric. An available host is discovered at a multi-pathing driver of the first host. The available host is capable of delivering I/Os to the primary array controller. An I/O is redirected from said first host to the available host over a secondary communication network for delivery to the storage system.