Shared Storage Failure Recovery via Central Configuration Repository

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In a shared disk storage system, the loss of communication links between nodes can lead to a 'split brain' condition, where data corruption risks increase due to uncoordinated changes in storage configuration, as nodes cannot manage shared storage resources collectively.

Innovation Solution

A system and method where each node has two ports coupled to storage enclosures, allowing configuration information to be saved to a central location in shared storage, enabling nodes to communicate configuration changes even when conventional links fail, using a reservation system to prevent data corruption by ensuring nodes acknowledge changes before accessing storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If nodes use conventional communication links to manage shared storage resources, then coordination and data integrity are maintained, but system reliability deteriorates when links fail due to split brain conditions

Engineering Contradiction:
Improvesystem reliabilityVSAvoiddata corruption risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces shared storage as an intermediary medium for communication between nodes. When conventional communication links fail, nodes can still exchange configuration information through the shared storage system, acting as a mediator that maintains coordination without requiring direct node-to-node communication channels.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements a reservation system where nodes reserve the right to modify configuration information in shared storage before actually making changes. This preliminary reservation mechanism ensures that only one node can modify configuration at a time, preventing split brain conditions and data corruption even when communication links are degraded.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If nodes can access shared storage independently when communication links fail, then accessibility is maintained, but data integrity deteriorates due to uncoordinated changes

Engineering Contradiction:
Improvestorage accessibilityVSAvoiddata integrity
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements a feedback mechanism through the reservation system that notifies nodes of configuration changes made by other nodes. Before a node can access or modify shared storage, it must check the reservation status and acknowledge any pending changes, ensuring data integrity is maintained while still allowing independent access when communication links are functional.

Inventive Principle:
Principle #23Feedback

3Reliability

If a reservation system is implemented to prevent split brain conditions, then data integrity is improved, but system complexity increases

Engineering Contradiction:
Improvedata integrityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent makes the shared storage system multi-functional by using it both for data storage and for communication/coordination between nodes. The reservation system leverages existing shared storage resources to manage configuration information, eliminating the need for separate complex coordination hardware and reducing overall system complexity while maintaining data integrity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7577865B2System and method for failure recovery in a shared storage system
Publication Date: 2009.08.18 DELL PROD LP
  • US7577865B2 patent drawing
  • US7577865B2 patent drawing
  • US7577865B2 patent drawing

AI summary

A system and method is disclosed for failure recovery and communications in a shared storage system. The shared storage system includes at least two host nodes, each of which includes two ports. Each of the ports of each of the nodes is coupled to input ports of a storage enclosure. The input ports of the storage enclosures are in turn coupled to one another to form communications links between each of the host nodes. When the communications links between the host nodes fail, the host nodes are able to pass configuration information to each other by saving configuration information to a central location in a shared storage, such as a dedicated location in one of the storage drives of the storage enclosure that is directly coupled to both host nodes. The host nodes are able to force their peer nodes to read configuration changes before accessing possibly corrupted data from a previous configuration.