Fail-back Coordination for Storage Path Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed computing systems, multiple hosts sharing a storage array often lose synchronization regarding which paths are designated for use, leading to inconsistencies and performance degradation during fail-back operations, especially in active/passive arrays with auto-trespass features.
Innovation Solution
A system and procedure for coordinating fail-backs among multiple hosts, involving monitoring of data paths, polling nodes for access restoration, and centralized approval through a master host to ensure simultaneous and consistent resumption of primary paths across all hosts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If automatic fail-back is enabled in active/passive storage arrays with auto-trespass features, then path restoration speed is improved, but path assignment inconsistencies and performance degradation occur
Solution Approach 1:
The patent introduces a fail-back coordination mechanism that acts as an intermediary between multiple hosts and the storage array. This coordination mechanism manages the fail-back process centrally, ensuring that path assignments remain consistent across all hosts while maintaining fast restoration speeds. The coordinator prevents path assignment inconsistencies by controlling when and how fail-back operations occur.
2Loss of time
If hosts independently perform fail-back operations, then individual host recovery time is reduced, but system-wide path assignment mismatches increase
Solution Approach 1:
The patent implements a feedback-based coordination mechanism where hosts report their fail-back status to a central coordinator, which then manages the overall fail-back timing. This feedback loop ensures that all hosts remain synchronized in their path assignments while allowing rapid individual recovery. The coordinator receives status information from hosts and coordinates fail-back operations to prevent mismatches.
Data Source
AI summary
Systems and procedures may be used to coordinate the fail-back of multiple hosts in environments where the hosts share one or more data-storage resources. In one implementation, a procedure for coordinating fail-backs includes monitoring a failed data path to detect a restoration of the data path, polling remaining nodes in response to the restoration, and allowing the first node to resume communications if access has been restored to the remaining nodes.


