Standby HBAs for SCSI Target High Availability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current high availability solutions for SCSI targets over fibre channels in computing environments face limitations, including disruptive failures, dependency on specific tape drivers, and incomplete redundancy, leading to potential data loss and system instability.
Innovation Solution
The implementation of duplicate, standby host-bus adaptors (HBAs) with duplicate credentials for alternate nodes, allowing seamless failover by shutting down a failing node and transferring its jobs to an active node using a shared data structure, thereby ensuring continuous operation without the need for dedicated host drivers or I/O virtualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If duplicate standby HBAs with duplicate credentials are implemented for alternate nodes, then high availability and system reliability are improved, but device complexity increases
Solution Approach 1:
The patent implements duplicate standby HBAs with duplicate credentials (WWNN and WWPN) for alternate nodes. Each node maintains copies of the other node's HBA credentials, enabling seamless failover. When a node fails, the standby HBA on the alternate node can immediately assume the failed node's identity and take over its jobs without requiring driver reconfiguration or I/O virtualization, thus achieving high availability while managing complexity through standardized credential duplication.
2Productivity
If standby nodes take over jobs of failed nodes using shared data structures, then productivity and operational continuity are improved, but loss of time during failover occurs
Solution Approach 1:
The patent prepares standby HBAs in advance with duplicate credentials and configures shared data structures containing job information before failures occur. When a node fails, the standby HBA can immediately assume the failed node's identity and access pre-configured job information from shared data structures, minimizing failover time. The shared data structure enables rapid job transfer by providing direct access to SCSI device assignments, reservations, and host notification requirements without requiring time-consuming data reconstruction.
3Reliability
If fence devices are used to shut down failed nodes, then data integrity and system stability are improved, but device complexity increases
Solution Approach 1:
The patent introduces fence devices as intermediary components that mediate between the standby node and the failed node. When a node fails, the fence device on the standby node shuts down the failed node's HBAs to prevent split-brain scenarios and ensure data integrity. This intermediary approach maintains system stability and prevents data corruption while managing the complexity of coordinated shutdown procedures through standardized fencing mechanisms.
4Ease of operation
If duplicate credentials are used for standby HBAs, then ease of operation and seamless failover are improved, but loss of information during credential management may occur
Solution Approach 1:
The patent implements feedback mechanisms that monitor and verify the accuracy of duplicate credentials (WWNN and WWPN) between nodes. The system continuously checks that standby HBAs have correct duplicate credentials and that credential duplication maintains data integrity. This feedback approach ensures seamless failover by verifying credential accuracy while preventing information loss through continuous validation and error detection mechanisms.
Data Source
AI summary
For efficient high availability for a multi-node cluster using a processor device in a computing environment, using duplicate, standby host-bus adaptors (HBAs) for alternate nodes with respect to a node with the duplicate, standby HBAs using duplicate credentials of active HBAs of the node for shutting down the node, taking an active HBA of the node offline, and/or activating one of the alternate nodes. A failed node is shut down by an active one of the plurality of alternate nodes using a fence device, and all jobs of the failed node are taken over by the active node using a shared data structure including SCSI device assignments, SCSI reservations, and required host notifications.


