This invention proposes an automatic fault detection and
repair method and apparatus for GPU nodes. The method includes: real-time acquisition of hardware operation indicators,
software status indicators, and resource load indicators of
graphics processing unit nodes, and streaming the acquired indicator data to an analysis terminal; performing time-series analysis on the indicator data based on a
machine learning model to predict potential fault risks, and comparing the indicator data against thresholds using a preset rule base to identify faults that have occurred, thereby generating fault diagnosis results; matching corresponding repair strategies from a multi-level self-healing strategy
library according to the fault type and
severity level, and generating execution instructions; responding to the execution instructions, performing corresponding isolation marking, lossless migration of computing tasks,
system reset, or resource masking operations on the
graphics processing unit nodes, and verifying the node status after the operation is completed to determine whether to restore the node to the cluster or trigger a manual maintenance notification.