Datacenter Disaster Recovery Node Prioritization for Faster Outage Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Power outages during natural disasters can cause significant damage and disrupt business continuity in datacenters by requiring all nodes to be powered on simultaneously, affecting critical and non-critical nodes equally, which prolongs service restoration time.

Innovation Solution

A disaster recovery system that detects resource metrics, computes priority scores based on these metrics, predicts resource usage, and powers on nodes based on these scores and predictions to prioritize critical nodes for faster service restoration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If all nodes are powered on simultaneously, then service restoration is simplified, but service availability is reduced due to longer recovery time

Engineering Contradiction:
Improvepower restoration simplicityVSAvoidservice availability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments nodes into different priority groups (critical vs. non-critical) based on their role and importance to service functionality. Critical nodes are powered on first to restore essential services, while non-critical nodes are powered on later. This segmentation resolves the contradiction by enabling simplified simultaneous power restoration while improving service availability through staged activation of essential services first.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary assessment and classification of nodes before power restoration, determining priority scores based on node roles, service dependencies, and resource usage patterns. This preliminary action enables the system to pre-plan the power-on sequence, ensuring critical services are restored first without requiring complex real-time decision-making during the actual power restoration process.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If critical and non-critical nodes are treated equally, then node management is simplified, but service restoration time is prolonged

Engineering Contradiction:
Improvenode management complexityVSAvoidservice restoration time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent applies local quality by assigning different treatment levels to different nodes based on their specific characteristics and roles. Critical nodes receive prioritized power-on treatment while non-critical nodes follow standard procedures. This differentiation is implemented through priority scoring mechanisms that automatically distinguish between node types, reducing manual management complexity while significantly reducing service restoration time through targeted critical node activation.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts node power-on priorities based on real-time conditions, resource availability, and predicted service needs. The priority scoring mechanism continuously evaluates node importance and adapts the power restoration sequence accordingly. This dynamic approach simplifies management by automating decision-making while reducing restoration time through adaptive prioritization of critical services based on current system state.

Inventive Principle:
Principle #15Dynamics

3Productivity

If priority-based power-on is implemented, then service restoration speed is improved, but system complexity increases

Engineering Contradiction:
Improveservice restoration speedVSAvoidpower management system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling the power management system to automatically assess node priorities, calculate priority scores, and determine power-on sequences without human intervention. The system uses built-in monitoring of node roles, service dependencies, and resource usage patterns to autonomously make prioritization decisions. This self-service mechanism significantly improves service restoration speed while minimizing the increase in system complexity through automation of what would otherwise require manual configuration.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback mechanisms that continuously monitor node status, service availability, and resource usage during the power restoration process. This feedback information is used to adjust priority scores and refine the power-on sequence in real-time. The feedback loop enables the system to learn from actual system behavior and optimize future power restoration operations, improving productivity while keeping complexity manageable through data-driven automation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12547509B2Intelligent disaster recovery of datacenter cloud outage
Publication Date: 2026.02.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12547509B2 patent drawing
  • US12547509B2 patent drawing
  • US12547509B2 patent drawing

AI summary

An embodiment includes detecting by a resource module of a disaster recovery system of a datacenter a resource metric. The embodiment includes responsive to the resource metric, computing a priority score by a priority scoring module of the disaster recovery system based on the resource metric. The embodiment includes predicting, by a prediction module of the disaster recovery system, a predicted resource usage metric based in part on the resource metric. The embodiment also includes powering on a node of the datacenter by a power manager module of the disaster recovery system based on the priority score and the predicted resource usage metric.