Node anomaly detection method and device, electronic device, and storage medium

By performing risk prediction and multi-dimensional detection on computing cluster nodes, abnormal nodes can be quickly screened and isolated, solving the problems of insufficient predictability and silent data corruption in existing technologies, and improving the overall performance and availability of computing clusters.

CN122372410APending Publication Date: 2026-07-10MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610581691.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-28
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies struggle to predictively detect node anomalies when managing and maintaining ultra-large-scale computing clusters, leading to delayed fault location, performance degradation, and ineffective protection against silent data corruption, thus impacting the overall performance and availability of the computing cluster.

Method used

By performing risk prediction on nodes in the computing cluster, determining the predicted probability of various anomaly types, screening out high-risk nodes, and using multi-dimensional detection components to detect anomalies, abnormal nodes can be quickly located, isolated, or repaired.

Benefits of technology

It improved the training success rate and overall performance of the computing cluster, increased the availability of computing nodes, and maintained high performance and high availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122372410A_ABST
    Figure CN122372410A_ABST
Patent Text Reader

Abstract

This disclosure provides a node anomaly detection method, apparatus, electronic device, and storage medium. The method is applied to a computing cluster comprising multiple computing nodes. The method includes: performing risk prediction on a target node to be predicted among the multiple computing nodes; determining the predicted anomaly probability of the target node for multiple anomaly types, including multiple types such as faults, slow nodes, and silent data corruption (SDC); identifying risk nodes from the target nodes based on the predicted anomaly probabilities; performing anomaly detection on the risk nodes based on the predicted anomaly probability of each risk node to obtain anomaly detection results for the risk nodes; and identifying anomaly nodes from the risk nodes based on the anomaly detection results. Embodiments of this disclosure can improve the overall performance of the computing cluster and the proportion of computing nodes in the cluster that are in an available state.
Need to check novelty before this filing date? Find Prior Art