MPIO Driver Failure Detection and Path Update

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In information processing systems, delays in detecting storage system port failures and restorations across multiple host devices can lead to sub-optimal performance in load balancing and failover policy execution due to inconsistent and delayed notifications.

Innovation Solution

Implementing a multi-path layer with MPIO drivers that detect target failure status and update path availability by obtaining and processing failure status information from storage systems, ensuring all host devices learn of port failures and restorations quickly and efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If traditional RSCN notification mechanisms are used for port failure detection, then host devices with active IO operations can learn of failures relatively quickly, but host devices without active IO operations experience substantial delays in learning of failures and restorations

Engineering Contradiction:
Improvefailure detection speedVSAvoidconsistent failure notification across all host devices
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent merges the failure detection mechanisms of all host devices through a centralized storage system that actively pings targets and distributes failure status information to all hosts. This combines the active detection capability of hosts with IO operations with the passive detection capability of hosts without active operations, ensuring all hosts receive consistent failure notifications simultaneously.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The storage system implements a feedback mechanism where it actively pings targets, detects failures, and notifies all host devices of the failure status. This closed-loop feedback ensures that all hosts receive timely and consistent failure information regardless of their current IO operation state, resolving the inconsistency in failure detection timing.

Inventive Principle:
Principle #23Feedback

2Productivity

If host devices rely on active IO operations to detect target failures, then detection can occur quickly for those devices, but load balancing and failover policies are adversely impacted due to delayed detection on other devices

Engineering Contradiction:
Improveload balancing efficiencyVSAvoiddelay in learning of port restoration
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The storage system performs preliminary failure detection by actively pinging targets before hosts need to detect failures through IO operations. This preliminary action ensures that failure status is known in advance and can be immediately communicated to all host devices, enabling timely load balancing and failover decisions without waiting for IO operation-based detection.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If registered state change notifications (RSCNs) are used to notify host devices of port failures, then some host devices can learn of failures quickly, but notification reliability is compromised as messages may not be received by all host devices or received only after substantial delay

Engineering Contradiction:
Improvefailure notification deliveryVSAvoiddelay in failure notification receipt
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The storage system acts as an intermediary that centralizes failure detection and notification distribution. Instead of relying on distributed RSCN messages that may be lost or delayed, the storage system directly pings targets, determines failure status, and notifies all host devices through a controlled mechanism, ensuring reliable and timely delivery of failure information to all hosts.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11012510B2Host device with multi-path layer configured for detecting target failure status and updating path availability
Publication Date: 2021.05.18 EMC IP HLDG CO LLC
  • US11012510B2 patent drawing
  • US11012510B2 patent drawing
  • US11012510B2 patent drawing

AI summary

A host device is configured to communicate over a network with a storage system comprising a plurality of storage devices. The host device comprises a multi-path input-output (MPIO) driver configured to control delivery of input-output (IO) operations from the host device to the storage system over selected ones of a plurality of paths through the network, where the paths are associated with respective initiator-target pairs, and each of a plurality of targets of the initiator-target pairs comprises a corresponding port of the storage system. The MPIO driver is further configured to obtain from the storage system information characterizing failure status of at least a subset of the targets, and to update availability status of the paths based at least in part on the obtained information.