Robustness Scoring for Computing Infrastructure Nodes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computing infrastructures face challenges in predicting and preventing downtime in disaggregated environments, where heterogeneous nodes and complex web connections lead to increased complexity and reactive maintenance, resulting in potential financial losses and reliability issues.

Innovation Solution

A robustness scoring system that interprets metadata and telemetry data to generate health scores for nodes and hardware components, predicting potential failures and enabling proactive mitigation and self-healing by correlating error cases and learning from component behavior.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If disaggregated computing environments with heterogeneous nodes are used, then computing power and flexibility are improved, but system complexity and difficulty of predicting failures increase

Engineering Contradiction:
Improvecomputing power and flexibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the computing infrastructure into discrete nodes and further into individual hardware components (CPU, GPU, storage, memory, network interfaces). Each component is independently monitored and assessed for robustness, allowing the system to manage complexity through modular evaluation rather than treating the entire heterogeneous system as a monolithic unit.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transforms the complex qualitative assessment of system health into quantitative robustness scores (0-1 scale) derived from multiple telemetry parameters. By changing the state of monitoring from binary (healthy/unhealthy) to continuous (robustness score), the system can effectively handle the complexity of heterogeneous environments with standardized metrics.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If comprehensive telemetry data collection is implemented, then failure prediction accuracy is improved, but data processing complexity and resource consumption increase

Engineering Contradiction:
Improvefailure prediction accuracyVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts only the most relevant telemetry parameters needed for robustness assessment from the vast amount of available system data. Instead of processing all telemetry data, the system selectively collects and processes specific parameters (error rates, performance metrics, utilization levels) that directly correlate with component failure risk, reducing processing complexity while maintaining prediction accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The robustness score acts as an intermediary that simplifies complex telemetry data into a single interpretable metric. Rather than directly analyzing multiple complex parameters, the system uses the robustness score as an intermediate representation that encapsulates the state of hardware components, making failure prediction more manageable.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If reactive maintenance is used, then immediate response to failures is achieved, but unplanned downtime and financial losses increase

Engineering Contradiction:
Improveresponse speedVSAvoidup-time
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system performs preliminary actions by continuously monitoring hardware robustness scores and identifying components at risk of failure before actual failures occur. By detecting declining robustness trends and predicting potential failures in advance, the system enables proactive maintenance scheduling, workload migration, and resource reallocation before downtime occurs, thus improving up-time while maintaining operational responsiveness.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11693721B2Creating robustness scores for selected portions of a computing infrastructure
Publication Date: 2023.07.04 INTEL CORP
  • US11693721B2 patent drawing
  • US11693721B2 patent drawing
  • US11693721B2 patent drawing

AI summary

A system for generating a robustness score for hardware components, nodes, and clusters of nodes in a computing infrastructure is provided. The system includes a memory and at least one processing device coupled to the memory. The processing device is to obtain first telemetry data associated with a selected portion of a computing infrastructure, and the selected portion includes a first node and a first hardware component. The processing device is further to obtain first metadata associated with the selected portion, input one or more telemetry inputs corresponding to the first telemetry data into a machine learning model, input one or more metadata inputs corresponding to the first metadata into the machine learning model, and generate, from the machine learning model, a first robustness score for the first hardware component representing a health state of the first hardware component.