Multi-Node Power Control for Automatic Recovery After Partial Outages

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-node systems, partial power outages can lead to inactive nodes, causing the remaining active nodes to become blocked, necessitating manual intervention and prolonged downtime for I/O processing to resume.

Innovation Solution

A multi-node system with controllers, volatile and nonvolatile memories, and power supply control devices that detect inactive nodes, save necessary data, and restart processing upon power recovery, enabling automatic resumption of I/O operations without manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If the Auto-PSON function is implemented to automatically turn on power upon recovery from power outage, then all nodes can start up to reach the state where I/O processing can be performed, but in case of partial outage, the remaining nodes stay blocked and the system cannot return to the state before the partial outage

Engineering Contradiction:
Improveautomatic startup capabilityVSAvoidsystem operational status after partial outage
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The system performs preliminary actions by detecting inactive nodes and saving necessary data from volatile memory to nonvolatile memory before the power supply control device restarts the processor. This preliminary data preservation ensures that when power is restored, the system can resume operations without requiring manual intervention, thus resolving the contradiction between automatic startup and system reliability after partial outage.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If maintenance personnel manually restart the remaining nodes after partial outage, then the system can return to normal operation, but it takes an extended period of time for the multi-node system to resume I/O processing

Engineering Contradiction:
Improvesystem recovery capabilityVSAvoiddowntime for I/O processing resumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements self-service by automatically detecting inactive nodes, determining whether system operation can continue, saving necessary data to nonvolatile memory, and restarting the processor through the power supply control device. This automated self-recovery process eliminates the need for manual intervention by maintenance personnel and significantly reduces the time required to resume I/O processing, thus resolving the contradiction between system recovery capability and downtime.

Inventive Principle:
Principle #25Self-service

3Productivity

If the system saves necessary data from volatile memory to nonvolatile memory upon detecting inability to continue operation, then the system can automatically resume processing, but this requires additional memory operations and processing steps

Engineering Contradiction:
Improvesystem recovery speedVSAvoiddata saving and restart mechanism
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system extracts and preserves only the necessary data from volatile memory to nonvolatile memory based on the determination of whether system operation can continue. This selective data extraction approach minimizes the amount of data that needs to be saved and processed, reducing the complexity of the data saving mechanism while maintaining high productivity during system recovery.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12393247B2Multi-node system and power supply control method
Publication Date: 2025.08.19 HITACHI VANTARA LTD
  • US12393247B2 patent drawing
  • US12393247B2 patent drawing
  • US12393247B2 patent drawing

AI summary

In the event of a partial outage, a multi-node system is enabled to start processing easily and appropriately. The multi-node system includes multiple nodes each including at least one controller, the controller including a processor, a power supply control microcomputer, a memory, and a nonvolatile memory. The processor detects whether or not any one of the nodes is inactive due to a power outage. The processor determines whether or not operation of the multi-node system can be continued, on the basis of operational status of the nodes. Upon determination that the operation of the multi-node system cannot be continued, the processor saves necessary data held in the memory into the nonvolatile memory. The power supply control microcomputer restarts the processor. When the node in the power outage has recovered therefrom following the restart, the multi-node system is caused to start processing.