Preemptive Resource Diagnostics in Disaggregated Data Centers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional computing systems face challenges in performing deep health diagnostics on resources without disrupting critical workloads, leading to potential resource failures and increased costs due to the inability to interrupt or replace resources while they are in use.

Innovation Solution

A disaggregated computing environment allows for dynamic resource allocation and health check diagnostics to be performed without disrupting workloads, by identifying failure patterns, predicting resource failures, and swapping resources in real-time, enabling proactive replacement and optimization of resource usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep health diagnostics are performed on resources, then resource failure prediction accuracy is improved, but workload disruption occurs

Engineering Contradiction:
Improveresource failure prediction accuracyVSAvoidworkload execution continuity
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system segments resources into multiple interchangeable units within a disaggregated architecture. When deep diagnostics are needed, only the specific resource undergoing testing is temporarily removed from its workload, while other resources continue operating. This segmentation allows diagnostics on individual components without system-wide disruption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediary components including a resource manager and virtualization layer that mediate between workloads and physical resources. When a resource requires deep diagnostics, the intermediary layer redirects workloads to alternative resources or virtual representations, allowing the original resource to be tested without directly impacting workload execution.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If resources are replaced proactively to prevent failures, then system reliability is improved, but resource allocation complexity increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoidresource allocation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary health assessments and deep diagnostics on resources before actual failures occur. By continuously monitoring health metrics and conducting predictive analytics, the system identifies resources at risk and proactively replaces them with healthy alternatives from the resource pool, preventing failures before they impact workloads.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements continuous feedback loops through health monitoring systems that track resource performance and predict failures. This feedback mechanism provides real-time information to the resource manager, enabling dynamic allocation decisions. The feedback loop continuously refines failure predictions and optimizes resource replacement timing, managing complexity through automated decision-making.

Inventive Principle:
Principle #23Feedback

3Reliability

If resources are swapped in real-time without interruption, then service availability is improved, but diagnostic depth is reduced

Engineering Contradiction:
Improveservice availabilityVSAvoiddiagnostic depth
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system dynamically adjusts resource allocation based on real-time health assessments. When resources show signs of degradation, the system can swap them out during low-utilization periods or redirect workloads, maintaining service availability. The diagnostic process is dynamic, adapting test intensity and duration based on resource criticality and current workload conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary health checks continuously in the background before failures occur. When resources show early signs of problems, deep diagnostics are initiated proactively during scheduled maintenance windows or low-utilization periods. This preliminary action allows comprehensive testing without interrupting critical workloads, as healthy replacement resources are already prepared in the resource pool.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10761915B2Preemptive deep diagnostics and health checking of resources in disaggregated data centers
Publication Date: 2020.09.01 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10761915B2 patent drawing
  • US10761915B2 patent drawing
  • US10761915B2 patent drawing

AI summary

Embodiments for preemptive deep diagnostics of resources in a disaggregated computing environment. Responsive to detecting a threshold breach of a recurrent event associated with a first resource of a first resource type executing a workload, an alert is generated; and responsive to receiving the alert, the execution of the workload on the first resource is ceased. Health check diagnostics are identified and invoked on the first resource based on the alert and a server telemetry. Results of the health check diagnostics are mapped to a set of learned failure patterns; and a potential failure of the first resource is predicted based on the mapping.