Storage Cluster Communications Session Failover via Preliminary Action
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Remote storage cluster systems face challenges in maintaining communication sessions and ensuring fault tolerance due to complex architectures and limitations in client device protocols, which hinder effective handling of errors and data access command failures.
Innovation Solution
The system groups nodes into high availability (HA) groups geographically dispersed and interconnected via networks, with active and inactive communications sessions to enable seamless takeover and metadata duplication for quick reestablishment of sessions, ensuring continuous data access and fault tolerance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If nodes are geographically dispersed to enable fault tolerance, then system reliability is improved, but communication session maintenance becomes more complex
Solution Approach 1:
The system pre-establishes multiple inactive communications sessions between nodes before failures occur. When a failure is detected, these pre-configured sessions can be rapidly activated without requiring complex real-time negotiation, thus maintaining reliability while simplifying failure recovery procedures
Solution Approach 2:
The system creates duplicate metadata copies at multiple nodes and maintains multiple communication paths. When a node fails, the system can switch to a copy held by another node through a pre-established session, enabling fault tolerance without complex reconstruction procedures
2Adaptability or versatility
If multiple interconnected modules are used to perform storage tasks, then system capability is improved, but error handling complexity increases
Solution Approach 1:
The system divides error handling into distinct segments: detection (through monitoring communications sessions), isolation (by identifying failed nodes), and recovery (through pre-established alternative sessions). This modular approach manages complexity while maintaining versatile storage capabilities across multiple nodes
3Productivity
If rapid error recovery is implemented, then system availability is improved, but communication reestablishment time may increase
Solution Approach 1:
The system performs preliminary configuration of multiple communication sessions and metadata copies before failures occur. When failures are detected, the system can immediately switch to pre-configured alternatives, achieving rapid recovery without the delays of real-time negotiation or reconstruction
Solution Approach 2:
The system continuously monitors communications sessions to detect failures early. This feedback mechanism enables the system to initiate recovery procedures immediately upon detecting issues, minimizing downtime while using pre-established paths to avoid reestablishment delays
Data Source
AI summary
Various embodiments are generally directed to techniques for preparing to respond to failures in performing a data access command to modify client device data in a storage cluster system. An apparatus may include a processor component of a first node coupled to a first storage device; an access component to perform a command on the first storage device; a replication component to exchange a replica of the command with the second node via a communications session formed between the first and second nodes to enable at least a partially parallel performance of the command by the first and second nodes; and a multipath component to change a state of the communications session from inactive to active to enable the exchange of the replica based on an indication of a failure within a third node that precludes performance of the command by the third node. Other embodiments are described and claimed.


