Fault Detection System for HPC Workload Scheduling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High-performance computing (HPC) systems face challenges in detecting faults during the execution of data-intensive workloads like HPC and AI workloads, which can lead to system failures and impact overall performance.

Innovation Solution

A computer-implemented method and system that improves workload scheduling by identifying faults in systems. This involves receiving a configuration file for a set of systems, retrieving event data from system executions, determining parameters from this data, and assigning output labels to workload termination parameters based on system status and configuration data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If fault detection is implemented in HPC systems, then system reliability is improved, but device complexity increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary fault detection system that acts as a mediator between the HPC workload and the underlying hardware systems. This intermediary component collects, processes, and analyzes fault information from multiple sources, thereby improving system reliability without directly increasing the complexity of the core computing systems. The intermediary layer consolidates complexity into a dedicated management function.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback mechanisms where fault detection information is continuously monitored and fed back into the workload management system. This feedback loop enables dynamic adjustments in workload scheduling and resource allocation based on real-time fault conditions, improving reliability while keeping the detection system's complexity manageable through automated responses.

Inventive Principle:
Principle #23Feedback

2Difficulty of detecting and measuring

If comprehensive fault monitoring is implemented, then fault detection capability is improved, but loss of time increases

Engineering Contradiction:
Improvefault detection capabilityVSAvoidtime for fault analysis
Core Design Contradiction:
Difficulty of detecting and measuringVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-configuring monitoring parameters, thresholds, and detection rules before faults occur. The system maintains pre-established frameworks for fault detection that enable rapid identification when anomalies arise, reducing the time needed for analysis during actual fault conditions while maintaining comprehensive monitoring capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs partial action by focusing fault monitoring on critical systems and parameters rather than attempting to monitor every possible component equally. This selective monitoring approach maintains adequate fault detection capability for the most important elements while significantly reducing the time and resources required for comprehensive analysis.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If workload scheduling optimization is implemented, then productivity is improved, but device complexity increases

Engineering Contradiction:
Improveworkload scheduling efficiencyVSAvoidscheduling system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements dynamics in the scheduling system by enabling adaptive, real-time adjustments to workload distribution based on detected fault conditions. The scheduling algorithm dynamically reassigns tasks to healthy systems when faults are detected, improving productivity through flexible response while managing complexity through rule-based automation rather than overly complex decision-making frameworks.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250181410A1System and method to improve scheduling of workload by identifying faults in systems
Publication Date: 2025.06.05 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250181410A1 patent drawing
  • US20250181410A1 patent drawing
  • US20250181410A1 patent drawing

AI summary

A method, system, and a computer program product for improving scheduling of workload by identifying faults in systems is disclosed. The present invention may include receiving a configuration file associated with a configuration of the set of systems used for execution of a workload. The present invention may include retrieving event data associated with a set of events. The present invention may include determining a first set of parameters from first event data associated with a first event of the set of events. The present invention may include assigning an output label to a first workload termination parameter based on the first set of parameters and a presence of first configuration data within the received configuration file. The present invention may include storing the assigned output label to the first workload termination parameter in memory.