Multi-pathing Driver I/O Redirection for Storage Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing failover protocols in enterprise storage arrays often result in costly and unnecessary downtime when connectivity is compromised, as they trigger cluster-wide failovers even for isolated node failures, leading to paused I/O processing and performance delays.
Innovation Solution
Implementing a method to redirect I/Os from a local host to an available remote host capable of delivering I/Os to the storage system, thereby avoiding the need for a cluster-wide failover protocol and maintaining access during connectivity failures or maintenance operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a failover protocol is implemented to ensure continuous access to storage arrays, then reliability is improved, but productivity deteriorates due to costly downtime and paused I/O processing
Solution Approach 1:
The patent segments the failover detection and response mechanism by node. Instead of a centralized failover protocol that affects the entire cluster, each node independently monitors its own connectivity to the storage array and executes failover only when necessary. This segmentation prevents cluster-wide I/O pauses when a single node experiences connectivity issues, thereby maintaining productivity while ensuring reliability for affected nodes.
Solution Approach 2:
The patent applies local quality by allowing different nodes in the cluster to have different failover states simultaneously. When one node experiences connectivity loss, only that specific node initiates failover to alternative paths, while other nodes continue normal I/O operations. This localized response maintains overall cluster productivity while ensuring continuous access for the affected node.
2Reliability
If a cluster-wide failover protocol is triggered to handle connectivity failures, then reliability is improved, but loss of time increases due to paused I/O processing across all nodes
Solution Approach 1:
The patent divides the failover execution into node-specific operations rather than a synchronized cluster-wide event. Each node independently determines when to pause I/O and when to switch paths, minimizing the time that any given node's I/O is paused. Other nodes experience no interruption, significantly reducing overall loss of time while maintaining reliability for the failed node.
Solution Approach 2:
The patent implements preliminary action by pre-configuring multiple active paths to the storage array before failures occur. When connectivity is lost on one node, the failover to pre-configured alternative paths can occur rapidly without requiring extensive reconfiguration or coordination with other nodes, thereby reducing the time I/O processing is paused.
3Adaptability or versatility
If maintenance operations are performed on a primary host controller, then adaptability is improved, but productivity deteriorates due to cluster-wide access prevention
Solution Approach 1:
The patent segments the maintenance operation impact to the specific node undergoing maintenance. When a primary host controller on one node is taken down for maintenance, only that node experiences connectivity loss and initiates failover. Other nodes in the cluster continue to access the storage array normally through their own primary controllers, maintaining productivity while allowing necessary maintenance to proceed on the affected node.
Solution Approach 2:
The patent applies local quality by allowing different nodes to operate in different states simultaneously. The node undergoing maintenance experiences failover to alternative paths, while other nodes maintain normal operation. This localized impact enables maintenance operations to proceed without sacrificing cluster-wide productivity.
Data Source
AI summary
A method and system for load balancing. The method includes determining that connectivity between a first host and a primary array controller of a storage system has failed. The first host is configured to send input/output messages (I/Os) to a storage system through a storage network fabric. An available host is discovered at a multi-pathing driver of the first host. The available host is capable of delivering I/Os to the primary array controller. An I/O is redirected from said first host to the available host over a secondary communication network for delivery to the storage system.


