Cluster Failover via Priority-Based Reset Delay Mechanism

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current high availability computer systems face challenges in efficient failover processes, particularly during network splits, where multiple standby systems may inadvertently access shared resources, leading to double access and increased failover times, which can delay system recovery and reduce availability.

Innovation Solution

Implementing a reset priority index and timed reset delay mechanism across active and standby computers, where each computer sets a unique reset delay time based on its priority, ensuring that only the highest priority system resets a malfunctioning system first, and communicates this reset to others to prevent multiple resets and facilitate smooth resource transfer.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all standby systems reset a malfunctioning system simultaneously during network split, then failover can be performed, but double access to shared resources occurs and system reliability deteriorates

Engineering Contradiction:
Improvesystem reliabilityVSAvoiddouble access to shared resources
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent assigns different reset priorities to different standby systems, creating local differentiation in their reset behavior. The highest priority standby system resets the malfunctioning system first, while lower priority systems wait, preventing simultaneous resets and double access to shared resources.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements a preliminary priority assignment mechanism before failover occurs. Standby systems are pre-configured with reset priorities, and the highest priority system is predetermined to perform the reset action first, preventing the harmful effect of simultaneous resets before it can occur.

Inventive Principle:
Principle #10Preliminary action

2Object-generated harmful factors

If standby systems wait for fixed time before resetting, then double access is prevented, but failover time increases and productivity decreases

Engineering Contradiction:
Improvedouble access preventionVSAvoidfailover time
Core Design Contradiction:
Object-generated harmful factorsVSProductivity

Solution Approach 1:

The patent replaces the static fixed waiting time approach with a dynamic priority-based timing mechanism. Instead of all systems waiting for the same duration, each standby system determines its reset timing based on its assigned priority, allowing the highest priority system to act immediately while others wait appropriately, thus preventing double access without unnecessary delays.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of reset timing from a uniform fixed value to a variable determined by priority level. This parameter change allows different standby systems to have different reset delays, optimizing both the prevention of double access and the speed of failover by eliminating unnecessary waiting time for high-priority systems.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If multiple standby systems become active simultaneously, then system availability is maintained, but resource contention occurs and reliability worsens

Engineering Contradiction:
Improvesystem availabilityVSAvoidresource contention
Core Design Contradiction:
ReliabilityVSLoss of substance

Solution Approach 1:

The patent introduces local differentiation among standby systems through priority assignment. Only the highest priority standby system is permitted to become active and access shared resources after a reset, while lower priority systems remain inactive, thereby preventing resource contention while maintaining system availability through controlled failover.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS7418627B2Cluster system wherein failover reset signals are sent from nodes according to their priority
Publication Date: 2008.08.26 HITACHI LTD
  • US7418627B2 patent drawing
  • US7418627B2 patent drawing
  • US7418627B2 patent drawing

AI summary

A high availability cluster computer system can realize exclusive control of a resource shared between computers and effect failover by resetting a currently-active system computer in case a malfunction occurs in the currently-active system computer. In case a malfunction occurs in a certain system in a cluster, another system in the cluster which has detected the malfunction issues a reset based on a priority to realize failover, in which a standby system takes over the processing of the malfunctioning system when the malfunctioning system is stopped.