Cluster Management Module Automatic Failover and Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cluster systems, such as those described in Japanese Patent Application Publication No. 2019-53587, are unable to handle failures without manual intervention, which can lead to system downtime and operational disruptions.
Innovation Solution
A cluster system with a management module configuration that includes a representative management module and a standby management module, equipped with failure monitoring, failover control, and recovery units, allowing for automatic failover and recovery without manual intervention. Each management module monitors for failures, switches roles when necessary, and restores the failure monitoring and failover units to ensure continuous operation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single representative management module is used to manage the cluster system, then device complexity is reduced, but reliability deteriorates because the system cannot handle failures without manual intervention
Solution Approach 1:
The standby management module automatically detects failures of the representative management module and performs failover without manual intervention. The system monitors itself and self-restores by switching to the standby module when a failure is detected, eliminating the need for external manual intervention while maintaining a relatively simple management structure.
Solution Approach 2:
A standby management module is pre-configured in the system before any failure occurs. This standby module remains ready to take over management functions immediately when the representative management module fails, ensuring continuous system operation without downtime or manual intervention.
2Reliability
If external management apparatuses are added to handle failures, then reliability is improved, but device complexity and hardware costs increase
Solution Approach 1:
The standby management module serves multiple functions: it monitors the representative management module's status, automatically detects failures, executes failover procedures, and can restore the representative module after recovery. This multi-functional component eliminates the need for separate external management apparatuses, maintaining reliability while avoiding additional hardware complexity.
Solution Approach 2:
The failure monitoring, failover control, and restoration functions are merged into the existing management modules rather than being implemented as separate external apparatuses. The standby management module integrates multiple responsibilities that would otherwise require separate systems, reducing overall device complexity while ensuring reliable failure handling.
Data Source
AI summary
A cluster system including a plurality of nodes, a plurality of clusters included in each node and a management module managing the cluster system and an arithmetic module, which are included in each of the clusters, wherein, among all the management modules included in the cluster system, one management module is set representative management module, in the individual clusters, one is set as a master management module, and another is set as a standby management module. Each of the management modules includes a failure monitoring unit and a failover control unit. When a failure in the representative management module is detected by any of the failure monitoring units, any of the management modules included in the non-representative management modules, is set as a new representative management module. A recovery unit restores the failure monitoring unit and the failover control unit in the management module in which a failure is detected.


