Shared Storage Failure Recovery via Central Configuration Repository
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In a shared disk storage system, the loss of communication links between nodes can lead to a 'split brain' condition, where data corruption risks increase due to uncoordinated changes in storage configuration, as nodes cannot manage shared storage resources collectively.
Innovation Solution
A system and method where each node has two ports coupled to storage enclosures, allowing configuration information to be saved to a central location in shared storage, enabling nodes to communicate configuration changes even when conventional links fail, using a reservation system to prevent data corruption by ensuring nodes acknowledge changes before accessing storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If nodes use conventional communication links to manage shared storage resources, then coordination and data integrity are maintained, but system reliability deteriorates when links fail due to split brain conditions
Solution Approach 1:
The patent introduces shared storage as an intermediary medium for communication between nodes. When conventional communication links fail, nodes can still exchange configuration information through the shared storage system, acting as a mediator that maintains coordination without requiring direct node-to-node communication channels.
Solution Approach 2:
The patent implements a reservation system where nodes reserve the right to modify configuration information in shared storage before actually making changes. This preliminary reservation mechanism ensures that only one node can modify configuration at a time, preventing split brain conditions and data corruption even when communication links are degraded.
2Ease of operation
If nodes can access shared storage independently when communication links fail, then accessibility is maintained, but data integrity deteriorates due to uncoordinated changes
Solution Approach 1:
The patent implements a feedback mechanism through the reservation system that notifies nodes of configuration changes made by other nodes. Before a node can access or modify shared storage, it must check the reservation status and acknowledge any pending changes, ensuring data integrity is maintained while still allowing independent access when communication links are functional.
3Reliability
If a reservation system is implemented to prevent split brain conditions, then data integrity is improved, but system complexity increases
Solution Approach 1:
The patent makes the shared storage system multi-functional by using it both for data storage and for communication/coordination between nodes. The reservation system leverages existing shared storage resources to manage configuration information, eliminating the need for separate complex coordination hardware and reducing overall system complexity while maintaining data integrity.
Data Source
AI summary
A system and method is disclosed for failure recovery and communications in a shared storage system. The shared storage system includes at least two host nodes, each of which includes two ports. Each of the ports of each of the nodes is coupled to input ports of a storage enclosure. The input ports of the storage enclosures are in turn coupled to one another to form communications links between each of the host nodes. When the communications links between the host nodes fail, the host nodes are able to pass configuration information to each other by saving configuration information to a central location in a shared storage, such as a dedicated location in one of the storage drives of the storage enclosure that is directly coupled to both host nodes. The host nodes are able to force their peer nodes to read configuration changes before accessing possibly corrupted data from a previous configuration.


