Predictive GPU Malfunction Detection via Daemon Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current GPU monitoring technologies only alert users of malfunctions after they occur, disrupting normal business operations by requiring immediate GPU replacement or migration of programs, which can be disruptive and inefficient.

Innovation Solution

A method and apparatus that predict GPU malfunctions by installing a daemon program to collect and compare GPU status parameters with pre-configured mean status fault parameters using statistical models, allowing for proactive replacement or migration of programs before a malfunction occurs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If current GPU monitoring technology is used to detect faults, then fault detection capability is achieved, but business operations are disrupted due to immediate replacement or migration required after malfunction

Engineering Contradiction:
Improvefault detection capabilityVSAvoidbusiness operations continuity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by collecting and analyzing GPU status parameters (temperature, power consumption, usage duration) before actual malfunction occurs. The system builds prediction models using historical data from normal and faulty GPUs, enabling proactive identification of GPUs at risk of failure. This allows replacement or migration to be performed in advance, avoiding disruption to business operations while maintaining reliable fault detection capability

Inventive Principle:
Principle #10Preliminary action

2Reliability

If GPU replacement or program migration is performed immediately after malfunction detection, then system reliability is maintained, but operational efficiency decreases due to disruption

Engineering Contradiction:
Improvesystem reliabilityVSAvoidoperational disruption time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by predicting GPU failures before they occur using statistical models that analyze temperature, power consumption, and usage duration patterns. By identifying at-risk GPUs in advance, the system enables scheduled replacement or migration during maintenance windows rather than forcing immediate disruptive actions, thereby maintaining system reliability while minimizing operational disruption time

Inventive Principle:
Principle #10Preliminary action

3Productivity

If predictive maintenance is implemented using statistical models, then operational efficiency is improved by avoiding disruption, but system complexity increases due to data collection and analysis requirements

Engineering Contradiction:
Improveoperational efficiencyVSAvoidmonitoring system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements predictive maintenance by collecting GPU status parameters (temperature, power consumption, usage duration) and using statistical models to predict failures before they occur. This preliminary analysis enables proactive maintenance scheduling, improving operational efficiency by avoiding disruptive emergency replacements. The complexity is managed through systematic data collection frameworks and statistical prediction models that process multiple parameters to generate reliable failure predictions

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10031797B2Method and apparatus for predicting GPU malfunctions
Publication Date: 2018.07.24 CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
  • US10031797B2 patent drawing
  • US10031797B2 patent drawing
  • US10031797B2 patent drawing

AI summary

A method of predicting GPU malfunctions includes installing a daemon program at a GPU node, the daemon program periodically collecting GPU status parameters corresponding to the GPU node at a pre-determined time period. The method also includes obtaining the GPU status parameters from the GPU node and comparing the obtained GPU status parameters with mean status fault parameters to determine whether the GPU is to malfunction, where the mean status fault parameters are obtained by use of a pre-configured statistical model. Prior to a GPU enters a malfunction state, the GPU can be replaced, or the programs executing on the GPU can be migrated to other GPUs for execution, without affecting the normal business operations.