Robustness Scoring for Computing Infrastructure Nodes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computing infrastructures face challenges in predicting and preventing downtime in disaggregated environments, where heterogeneous nodes and complex web connections lead to increased complexity and reactive maintenance, resulting in potential financial losses and reliability issues.
Innovation Solution
A robustness scoring system that interprets metadata and telemetry data to generate health scores for nodes and hardware components, predicting potential failures and enabling proactive mitigation and self-healing by correlating error cases and learning from component behavior.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If disaggregated computing environments with heterogeneous nodes are used, then computing power and flexibility are improved, but system complexity and difficulty of predicting failures increase
Solution Approach 1:
The patent segments the computing infrastructure into discrete nodes and further into individual hardware components (CPU, GPU, storage, memory, network interfaces). Each component is independently monitored and assessed for robustness, allowing the system to manage complexity through modular evaluation rather than treating the entire heterogeneous system as a monolithic unit.
Solution Approach 2:
The system transforms the complex qualitative assessment of system health into quantitative robustness scores (0-1 scale) derived from multiple telemetry parameters. By changing the state of monitoring from binary (healthy/unhealthy) to continuous (robustness score), the system can effectively handle the complexity of heterogeneous environments with standardized metrics.
2Reliability
If comprehensive telemetry data collection is implemented, then failure prediction accuracy is improved, but data processing complexity and resource consumption increase
Solution Approach 1:
The patent extracts only the most relevant telemetry parameters needed for robustness assessment from the vast amount of available system data. Instead of processing all telemetry data, the system selectively collects and processes specific parameters (error rates, performance metrics, utilization levels) that directly correlate with component failure risk, reducing processing complexity while maintaining prediction accuracy.
Solution Approach 2:
The robustness score acts as an intermediary that simplifies complex telemetry data into a single interpretable metric. Rather than directly analyzing multiple complex parameters, the system uses the robustness score as an intermediate representation that encapsulates the state of hardware components, making failure prediction more manageable.
3Ease of operation
If reactive maintenance is used, then immediate response to failures is achieved, but unplanned downtime and financial losses increase
Solution Approach 1:
The system performs preliminary actions by continuously monitoring hardware robustness scores and identifying components at risk of failure before actual failures occur. By detecting declining robustness trends and predicting potential failures in advance, the system enables proactive maintenance scheduling, workload migration, and resource reallocation before downtime occurs, thus improving up-time while maintaining operational responsiveness.
Data Source
AI summary
A system for generating a robustness score for hardware components, nodes, and clusters of nodes in a computing infrastructure is provided. The system includes a memory and at least one processing device coupled to the memory. The processing device is to obtain first telemetry data associated with a selected portion of a computing infrastructure, and the selected portion includes a first node and a first hardware component. The processing device is further to obtain first metadata associated with the selected portion, input one or more telemetry inputs corresponding to the first telemetry data into a machine learning model, input one or more metadata inputs corresponding to the first metadata into the machine learning model, and generate, from the machine learning model, a first robustness score for the first hardware component representing a health state of the first hardware component.


