MPIO Driver Failure Detection and Path Update
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In information processing systems, delays in detecting storage system port failures and restorations across multiple host devices can lead to sub-optimal performance in load balancing and failover policy execution due to inconsistent and delayed notifications.
Innovation Solution
Implementing a multi-path layer with MPIO drivers that detect target failure status and update path availability by obtaining and processing failure status information from storage systems, ensuring all host devices learn of port failures and restorations quickly and efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If traditional RSCN notification mechanisms are used for port failure detection, then host devices with active IO operations can learn of failures relatively quickly, but host devices without active IO operations experience substantial delays in learning of failures and restorations
Solution Approach 1:
The patent merges the failure detection mechanisms of all host devices through a centralized storage system that actively pings targets and distributes failure status information to all hosts. This combines the active detection capability of hosts with IO operations with the passive detection capability of hosts without active operations, ensuring all hosts receive consistent failure notifications simultaneously.
Solution Approach 2:
The storage system implements a feedback mechanism where it actively pings targets, detects failures, and notifies all host devices of the failure status. This closed-loop feedback ensures that all hosts receive timely and consistent failure information regardless of their current IO operation state, resolving the inconsistency in failure detection timing.
2Productivity
If host devices rely on active IO operations to detect target failures, then detection can occur quickly for those devices, but load balancing and failover policies are adversely impacted due to delayed detection on other devices
Solution Approach 1:
The storage system performs preliminary failure detection by actively pinging targets before hosts need to detect failures through IO operations. This preliminary action ensures that failure status is known in advance and can be immediately communicated to all host devices, enabling timely load balancing and failover decisions without waiting for IO operation-based detection.
3Reliability
If registered state change notifications (RSCNs) are used to notify host devices of port failures, then some host devices can learn of failures quickly, but notification reliability is compromised as messages may not be received by all host devices or received only after substantial delay
Solution Approach 1:
The storage system acts as an intermediary that centralizes failure detection and notification distribution. Instead of relying on distributed RSCN messages that may be lost or delayed, the storage system directly pings targets, determines failure status, and notifies all host devices through a controlled mechanism, ensuring reliable and timely delivery of failure information to all hosts.
Data Source
AI summary
A host device is configured to communicate over a network with a storage system comprising a plurality of storage devices. The host device comprises a multi-path input-output (MPIO) driver configured to control delivery of input-output (IO) operations from the host device to the storage system over selected ones of a plurality of paths through the network, where the paths are associated with respective initiator-target pairs, and each of a plurality of targets of the initiator-target pairs comprises a corresponding port of the storage system. The MPIO driver is further configured to obtain from the storage system information characterizing failure status of at least a subset of the targets, and to update availability status of the paths based at least in part on the obtained information.


