Multipath Driver Cognitive Coordination for SAN Intermittent Errors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Storage area networks (SANs) face challenges in detecting and isolating recurring intermittent failures, which can lead to performance degradation and application failures due to elusive and transient nature of these faults, causing repeated recovery loops and false triggers.
Innovation Solution
A method and system for multipath driver cognitive coordination that detects recurring intermittent errors by determining if path recovery actions have been initiated, moving the path into a degraded state temporarily to prevent access, and subsequently allowing access once recovery is confirmed, enhancing communication between fiber channel protocol drivers and SCSI drivers to differentiate between real errors and those induced by recovery efforts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If path recovery actions are automatically initiated upon detecting intermittent errors, then system availability is improved, but performance degradation occurs due to repeated recovery loops and false triggers
Solution Approach 1:
The system performs preliminary analysis of error patterns before initiating recovery actions. By detecting recurring intermittent errors and analyzing their characteristics, the system determines whether recovery actions are truly necessary, preventing false triggers and repeated recovery loops that degrade performance
Solution Approach 2:
The system implements feedback mechanisms to monitor the effectiveness of recovery actions. By tracking whether recovery actions resolve the underlying issue or merely temporarily mask recurring problems, the system can adjust its behavior to avoid initiating unnecessary recovery operations that harm performance
2Reliability
If multiple redundant paths are implemented for fault tolerance, then system reliability is improved, but device complexity increases
Solution Approach 1:
The system merges the functionality of multiple redundant paths into a unified management framework. By coordinating path selection and failure recovery across multiple paths through centralized logic, the system maintains fault tolerance while reducing the operational complexity of managing redundant configurations
3Reliability
If path access is restricted during error conditions, then error propagation is prevented, but loss of time occurs due to temporary path unavailability
Solution Approach 1:
The system dynamically adjusts path access restrictions based on real-time error conditions. Rather than imposing static time-based restrictions, the system continuously monitors error patterns and adjusts path availability accordingly, restricting access only when necessary to contain errors while minimizing unnecessary downtime
Data Source
AI summary
An aspect includes detecting a recurring intermittent error in a path of a network in a system that includes at least one data transmission port configured for connection to at least one shared data storage device via a plurality of paths of the network. It is determined by a path control module (PCM) in the network, whether a path recovery action has been initiated by a fiber channel protocol driver in the network. In response to determining that the path recovery action has not been initiated, the data transmission port is prevented from accessing the path for a specified time period by moving the path into a degraded sub-state, and subsequent to the specified time period the data transmission port is provided access to the path. In response to determining that the path recovery action has been initiated, the data transmission port is provided access to the path.


