Cluster Failover via Priority-Based Reset Delay Mechanism
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current high availability computer systems face challenges in efficient failover processes, particularly during network splits, where multiple standby systems may inadvertently access shared resources, leading to double access and increased failover times, which can delay system recovery and reduce availability.
Innovation Solution
Implementing a reset priority index and timed reset delay mechanism across active and standby computers, where each computer sets a unique reset delay time based on its priority, ensuring that only the highest priority system resets a malfunctioning system first, and communicates this reset to others to prevent multiple resets and facilitate smooth resource transfer.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all standby systems reset a malfunctioning system simultaneously during network split, then failover can be performed, but double access to shared resources occurs and system reliability deteriorates
Solution Approach 1:
The patent assigns different reset priorities to different standby systems, creating local differentiation in their reset behavior. The highest priority standby system resets the malfunctioning system first, while lower priority systems wait, preventing simultaneous resets and double access to shared resources.
Solution Approach 2:
The patent implements a preliminary priority assignment mechanism before failover occurs. Standby systems are pre-configured with reset priorities, and the highest priority system is predetermined to perform the reset action first, preventing the harmful effect of simultaneous resets before it can occur.
2Object-generated harmful factors
If standby systems wait for fixed time before resetting, then double access is prevented, but failover time increases and productivity decreases
Solution Approach 1:
The patent replaces the static fixed waiting time approach with a dynamic priority-based timing mechanism. Instead of all systems waiting for the same duration, each standby system determines its reset timing based on its assigned priority, allowing the highest priority system to act immediately while others wait appropriately, thus preventing double access without unnecessary delays.
Solution Approach 2:
The patent changes the parameter of reset timing from a uniform fixed value to a variable determined by priority level. This parameter change allows different standby systems to have different reset delays, optimizing both the prevention of double access and the speed of failover by eliminating unnecessary waiting time for high-priority systems.
3Reliability
If multiple standby systems become active simultaneously, then system availability is maintained, but resource contention occurs and reliability worsens
Solution Approach 1:
The patent introduces local differentiation among standby systems through priority assignment. Only the highest priority standby system is permitted to become active and access shared resources after a reset, while lower priority systems remain inactive, thereby preventing resource contention while maintaining system availability through controlled failover.
Data Source
AI summary
A high availability cluster computer system can realize exclusive control of a resource shared between computers and effect failover by resetting a currently-active system computer in case a malfunction occurs in the currently-active system computer. In case a malfunction occurs in a certain system in a cluster, another system in the cluster which has detected the malfunction issues a reset based on a priority to realize failover, in which a standby system takes over the processing of the malfunctioning system when the malfunctioning system is stopped.


