Software Fencing for Cluster Nodes Without Hardware
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Recovery from a malfunctioning storage node in an active-active system configuration is challenging due to the risk of data corruption on shared memory and resources, especially in systems without inherent hardware fencing, which can be costly and complex.
Innovation Solution
Implementing a fencing scheme that allows storage nodes to communicate through various means, including storage drives and network connections, to detect malfunctioning nodes and enforce self-fencing actions, such as stopping access to critical resources and initiating a reboot, without requiring additional specialized hardware.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If inherent hardware fencing mechanisms are implemented, then data corruption risk is reduced, but system cost and complexity increase
Solution Approach 1:
The patent replaces inherent hardware fencing mechanisms with a software-based fencing scheme. The storage nodes use software protocols and communication mechanisms to detect malfunctioning nodes and enforce fencing actions, eliminating the need for specialized hardware fencing components while maintaining data protection capabilities
Solution Approach 2:
The patent creates a virtual fencing mechanism that replicates the functionality of hardware fencing through software. By using communication paths through storage drives and network connections, the system creates a software copy of hardware fencing behavior, allowing malfunctioning nodes to be fenced without physical hardware intervention
2Difficulty of detecting and measuring
If multiple communication paths are used, then detection capability is improved, but system complexity increases
Solution Approach 1:
The patent makes existing storage drives and network connections serve multiple functions. These communication paths are used both for normal data communication and for fencing detection, eliminating the need for dedicated detection hardware while improving malfunction detection capability through existing infrastructure
Solution Approach 2:
The storage nodes use their own communication infrastructure to detect malfunctions in other nodes. The system self-services by utilizing its existing communication paths for both operational and diagnostic purposes, avoiding additional complexity from separate detection mechanisms
Data Source
AI summary
Techniques for providing a fencing scheme for cluster systems without inherent hardware fencing. Storage nodes in an HA cluster communicate with one another over communication paths implemented through a variety of mechanisms, including drives and network connections. Each storage node in the HA cluster executes a fencing enforcer component operable to enforce actions or processes for initiating fencing of itself and/or initiating self-fencing of another storage node in the HA cluster determined to be malfunctioning. By providing for richer communications between storage nodes in an HA cluster and a richer set of actions or processes for fencing a malfunctioning storage node including self-fencing, the malfunctioning storage node can be made to exit the HA cluster in a more controlled fashion. In addition, a remaining storage node in the HA cluster can more safely take on the role of primary storage node with reduced risk of data corruption on shared resources.


