Fault Detection System for HPC Workload Scheduling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-performance computing (HPC) systems face challenges in detecting faults during the execution of data-intensive workloads like HPC and AI workloads, which can lead to system failures and impact overall performance.
Innovation Solution
A computer-implemented method and system that improves workload scheduling by identifying faults in systems. This involves receiving a configuration file for a set of systems, retrieving event data from system executions, determining parameters from this data, and assigning output labels to workload termination parameters based on system status and configuration data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If fault detection is implemented in HPC systems, then system reliability is improved, but device complexity increases
Solution Approach 1:
The patent introduces an intermediary fault detection system that acts as a mediator between the HPC workload and the underlying hardware systems. This intermediary component collects, processes, and analyzes fault information from multiple sources, thereby improving system reliability without directly increasing the complexity of the core computing systems. The intermediary layer consolidates complexity into a dedicated management function.
Solution Approach 2:
The patent implements feedback mechanisms where fault detection information is continuously monitored and fed back into the workload management system. This feedback loop enables dynamic adjustments in workload scheduling and resource allocation based on real-time fault conditions, improving reliability while keeping the detection system's complexity manageable through automated responses.
2Difficulty of detecting and measuring
If comprehensive fault monitoring is implemented, then fault detection capability is improved, but loss of time increases
Solution Approach 1:
The patent applies preliminary action by pre-configuring monitoring parameters, thresholds, and detection rules before faults occur. The system maintains pre-established frameworks for fault detection that enable rapid identification when anomalies arise, reducing the time needed for analysis during actual fault conditions while maintaining comprehensive monitoring capability.
Solution Approach 2:
The patent employs partial action by focusing fault monitoring on critical systems and parameters rather than attempting to monitor every possible component equally. This selective monitoring approach maintains adequate fault detection capability for the most important elements while significantly reducing the time and resources required for comprehensive analysis.
3Productivity
If workload scheduling optimization is implemented, then productivity is improved, but device complexity increases
Solution Approach 1:
The patent implements dynamics in the scheduling system by enabling adaptive, real-time adjustments to workload distribution based on detected fault conditions. The scheduling algorithm dynamically reassigns tasks to healthy systems when faults are detected, improving productivity through flexible response while managing complexity through rule-based automation rather than overly complex decision-making frameworks.
Data Source
AI summary
A method, system, and a computer program product for improving scheduling of workload by identifying faults in systems is disclosed. The present invention may include receiving a configuration file associated with a configuration of the set of systems used for execution of a workload. The present invention may include retrieving event data associated with a set of events. The present invention may include determining a first set of parameters from first event data associated with a first event of the set of events. The present invention may include assigning an output label to a first workload termination parameter based on the first set of parameters and a presence of first configuration data within the received configuration file. The present invention may include storing the assigned output label to the first workload termination parameter in memory.


