Storage Cluster Communications Session Failover via Preliminary Action

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Remote storage cluster systems face challenges in maintaining communication sessions and ensuring fault tolerance due to complex architectures and limitations in client device protocols, which hinder effective handling of errors and data access command failures.

Innovation Solution

The system groups nodes into high availability (HA) groups geographically dispersed and interconnected via networks, with active and inactive communications sessions to enable seamless takeover and metadata duplication for quick reestablishment of sessions, ensuring continuous data access and fault tolerance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If nodes are geographically dispersed to enable fault tolerance, then system reliability is improved, but communication session maintenance becomes more complex

Engineering Contradiction:
Improvefault toleranceVSAvoidcommunication session maintenance
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system pre-establishes multiple inactive communications sessions between nodes before failures occur. When a failure is detected, these pre-configured sessions can be rapidly activated without requiring complex real-time negotiation, thus maintaining reliability while simplifying failure recovery procedures

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates duplicate metadata copies at multiple nodes and maintains multiple communication paths. When a node fails, the system can switch to a copy held by another node through a pre-established session, enabling fault tolerance without complex reconstruction procedures

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If multiple interconnected modules are used to perform storage tasks, then system capability is improved, but error handling complexity increases

Engineering Contradiction:
Improvestorage task capabilityVSAvoiderror handling
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system divides error handling into distinct segments: detection (through monitoring communications sessions), isolation (by identifying failed nodes), and recovery (through pre-established alternative sessions). This modular approach manages complexity while maintaining versatile storage capabilities across multiple nodes

Inventive Principle:
Principle #1Segmentation

3Productivity

If rapid error recovery is implemented, then system availability is improved, but communication reestablishment time may increase

Engineering Contradiction:
Improveerror recovery speedVSAvoidsession reestablishment delay
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary configuration of multiple communication sessions and metadata copies before failures occur. When failures are detected, the system can immediately switch to pre-configured alternatives, achieving rapid recovery without the delays of real-time negotiation or reconstruction

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors communications sessions to detect failures early. This feedback mechanism enables the system to initiate recovery procedures immediately upon detecting issues, minimizing downtime while using pre-established paths to avoid reestablishment delays

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11016866B2Techniques for maintaining communications sessions among nodes in a storage cluster system
Publication Date: 2021.05.25 NETAPP INC
  • US11016866B2 patent drawing
  • US11016866B2 patent drawing
  • US11016866B2 patent drawing

AI summary

Various embodiments are generally directed to techniques for preparing to respond to failures in performing a data access command to modify client device data in a storage cluster system. An apparatus may include a processor component of a first node coupled to a first storage device; an access component to perform a command on the first storage device; a replication component to exchange a replica of the command with the second node via a communications session formed between the first and second nodes to enable at least a partially parallel performance of the command by the first and second nodes; and a multipath component to change a state of the communications session from inactive to active to enable the exchange of the replica based on an indication of a failure within a third node that precludes performance of the command by the third node. Other embodiments are described and claimed.