Preemptive Resource Replacement in Disaggregated Data Centers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional computing systems face challenges in performing deep diagnostics on resources without disrupting critical workloads, leading to potential catastrophic failures due to the inability to proactively identify and replace faulty resources in real-time, especially in disaggregated environments where resources are dynamically composed and interchanged.

Innovation Solution

A system that analyzes failure patterns and mitigation actions in a disaggregated computing environment, allowing for preemptive resource replacement by dynamically allocating healthy resources, isolating suspicious ones, and performing deep diagnostics without disrupting workloads, using a management module to track resource health and estimate time to failure.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If deep diagnostics are performed on resources, then resource health can be identified, but critical workloads are disrupted

Engineering Contradiction:
Improveresource health identificationVSAvoidworkload operation continuity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary failure pattern analysis and resource health assessment before critical failures occur. By continuously monitoring resources and predicting failures in advance, the system can proactively replace unhealthy resources before they disrupt workloads, thus maintaining both reliability and productivity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary management layer that decouples the diagnostic process from workload execution. This intermediary layer monitors resource health independently and orchestrates resource replacement without directly interfering with workload operations, allowing diagnostics to be performed without disrupting critical workloads.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If resources are replaced preemptively, then system reliability is improved, but resource management complexity increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoidresource management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements feedback mechanisms where resource health data, failure patterns, and replacement outcomes are continuously collected and analyzed. This feedback loop enables the system to automatically adjust replacement strategies, optimize resource allocation, and improve reliability while managing complexity through data-driven decision-making.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system enables self-service resource replacement by automatically detecting unhealthy resources, selecting appropriate replacements, and orchestrating the swap process without manual intervention. This automation reduces management complexity while maintaining high system reliability through consistent, rule-based resource replacement.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If continuous monitoring is performed, then failure prediction accuracy is improved, but system overhead increases

Engineering Contradiction:
Improvefailure prediction accuracyVSAvoidsystem overhead
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system applies partial monitoring by focusing diagnostic efforts on resources showing early signs of failure or those critical to workload performance. Instead of uniformly monitoring all resources at maximum intensity, the system dynamically adjusts monitoring depth based on risk assessment, improving failure prediction accuracy for critical resources while reducing overhead for stable resources.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11188408B2Preemptive resource replacement according to failure pattern analysis in disaggregated data centers
Publication Date: 2021.11.30 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11188408B2 patent drawing
  • US11188408B2 patent drawing
  • US11188408B2 patent drawing

AI summary

Embodiments for preemptive substitution of resources in a disaggregated computing environment. Failure patterns and mitigation actions are analyzed for specific failures of respective resources within the disaggregated computing environment. Responsive to determining a failure threshold has been reached for a first resource of a first type of the respective resources, a mitigation action is performed according to the analyzed failure patterns. A result of the mitigation action is determined and the result is used to improve the failure pattern analyzation.