Cluster Management Module Automatic Failover and Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing cluster systems, such as those described in Japanese Patent Application Publication No. 2019-53587, are unable to handle failures without manual intervention, which can lead to system downtime and operational disruptions.

Innovation Solution

A cluster system with a management module configuration that includes a representative management module and a standby management module, equipped with failure monitoring, failover control, and recovery units, allowing for automatic failover and recovery without manual intervention. Each management module monitors for failures, switches roles when necessary, and restores the failure monitoring and failover units to ensure continuous operation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single representative management module is used to manage the cluster system, then device complexity is reduced, but reliability deteriorates because the system cannot handle failures without manual intervention

Engineering Contradiction:
Improvemanagement structureVSAvoidfailure handling capability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The standby management module automatically detects failures of the representative management module and performs failover without manual intervention. The system monitors itself and self-restores by switching to the standby module when a failure is detected, eliminating the need for external manual intervention while maintaining a relatively simple management structure.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

A standby management module is pre-configured in the system before any failure occurs. This standby module remains ready to take over management functions immediately when the representative management module fails, ensuring continuous system operation without downtime or manual intervention.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If external management apparatuses are added to handle failures, then reliability is improved, but device complexity and hardware costs increase

Engineering Contradiction:
Improvefailure handling capabilityVSAvoidsystem configuration
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The standby management module serves multiple functions: it monitors the representative management module's status, automatically detects failures, executes failover procedures, and can restore the representative module after recovery. This multi-functional component eliminates the need for separate external management apparatuses, maintaining reliability while avoiding additional hardware complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The failure monitoring, failover control, and restoration functions are merged into the existing management modules rather than being implemented as separate external apparatuses. The standby management module integrates multiple responsibilities that would otherwise require separate systems, reducing overall device complexity while ensuring reliable failure handling.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11853175B2Cluster system and restoration method that performs failover control
Publication Date: 2023.12.26 HITACHI VANTARA LTD
  • US11853175B2 patent drawing
  • US11853175B2 patent drawing
  • US11853175B2 patent drawing

AI summary

A cluster system including a plurality of nodes, a plurality of clusters included in each node and a management module managing the cluster system and an arithmetic module, which are included in each of the clusters, wherein, among all the management modules included in the cluster system, one management module is set representative management module, in the individual clusters, one is set as a master management module, and another is set as a standby management module. Each of the management modules includes a failure monitoring unit and a failover control unit. When a failure in the representative management module is detected by any of the failure monitoring units, any of the management modules included in the non-representative management modules, is set as a new representative management module. A recovery unit restores the failure monitoring unit and the failover control unit in the management module in which a failure is detected.