Grey fault processing method and device, medium and program product

Through the graph neural network model and health scoring mechanism, gray faults in the server cluster are automatically detected and migrated, solving the problem of hidden performance degradation that is difficult to detect in existing technologies, achieving fast and accurate fault handling, and improving system reliability.

CN120631741AActive Publication Date: 2025-09-12INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511129420.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-12
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing fault detection methods mainly target faults that cause system shutdowns. They are unable to effectively detect and handle slow faults (gray faults) that implicitly reduce system performance, and rely on manual detection, which is inefficient.

Method used

A graph neural network model is used to detect server cluster performance data. Through single-node gray fault detection and health scoring mechanism, combined with the network link topology model, computing tasks are automatically identified and migrated to healthy nodes, reducing manual intervention.

Benefits of technology

It significantly improves the speed and accuracy of gray fault detection, reduces the need for manual detection, and improves the reliability and overall efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631741A_ABST
    Figure CN120631741A_ABST
Patent Text Reader

Abstract

The invention discloses a gray fault processing method and device, a medium and a program product, and relates to the technical field of detection, and the method comprises the steps: collecting the performance data of all server nodes; according to the performance data of all the server nodes and the graph neural network model, detecting the server nodes with gray faults in the server cluster; in response to the grey fault of any server node in the server cluster, verifying the grey fault of the first server node through a single-node grey fault detection algorithm; determining a target migration server node through a server node health scoring mechanism and a topology model of a network link in response to the fact that the gray fault verification of the first server node is correct; and migrating the calculation task of the first server node to the target migration server node. According to the method, the gray fault in the server cluster is detected, the gray fault detection speed and accuracy are remarkably improved, the requirement of manual detection is reduced, and the reliability of the whole system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of detection technology, and in particular to a gray fault processing method, device, medium and program product. Background Art

[0002] Currently, related fault handling methods mainly rely on manual or automated detection of hardware redundancy after fault modeling. However, these methods are mainly designed for faults that cause system shutdowns and are not very effective in detecting slow faults (gray faults) that implicitly reduce system performance. For example, hardware redundancy may cause "gray faults", that is, the system performance gradually deteriorates without being noticed. Manual detection requires a lot of human resources and is difficult to respond in real time. Summary of the Invention

[0003] The present application provides a gray fault processing method, device, medium and program product, the method includes collecting performance data of all server nodes; detecting server nodes with gray faults in the server cluster based on the performance data of all server nodes and a graph neural network model; in response to a gray fault occurring in any server node in the server cluster, setting any server as the first server node, and verifying the gray fault of the first server node through a single-node gray fault detection algorithm; in response to the gray fault verification of the first server node being correct, determining the target migration server node through a server node health scoring mechanism and a topological model of a network link; migrating the computing task of the first server node to the target migration server node. The present application significantly improves the speed and accuracy of gray fault detection in a server cluster, reduces the need for manual detection, and improves the reliability of the overall system.

[0004] This application provides a gray fault handling method, which is applied to a gray fault handling system. The system includes a server node cluster. The method includes: Collect performance data of all server nodes; Detect gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

[0005] This application also provides a gray fault handling device, including: Collection module, used to collect performance data of all server nodes; The first detection module is used to detect server nodes with gray faults in the server cluster based on the performance data of all server nodes and the graph neural network model; A second detection module is configured to, in response to a gray fault occurring on any server node in the server cluster, set any server as a first server node and verify the gray fault of the first server node using a single-node gray fault detection algorithm; a determination module configured to determine a target migration server node based on a server node health scoring mechanism and a topology model of a network link in response to a gray fault check of the first server node being correct; The processing module is configured to migrate the computing task of the first server node to the target migration server node.

[0006] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of a gray fault processing method when executing the computer program. The method comprises: Collect performance data of all server nodes; Detect gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

[0007] The present application further provides a computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the gray fault processing method are implemented. The method includes: Collect performance data of all server nodes; Detect gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

[0008] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the gray fault processing method are implemented. The method includes: Collect performance data of all server nodes; Detect gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

[0009] Through this application, since the method includes collecting performance data of all server nodes; detecting server nodes with gray faults in the server cluster based on the performance data of all server nodes and the graph neural network model; in response to a gray fault in any server node in the server cluster, setting any server as the first server node, and verifying the gray fault of the first server node through a single-node gray fault detection algorithm; in response to the gray fault verification of the first server node being correct, determining the target migration server node through the server node health scoring mechanism and the topology model of the network link; migrating the computing task of the first server node to the target migration server node. This application significantly improves the speed and accuracy of gray fault detection in server clusters, reduces the need for manual detection, and improves the reliability of the overall system.

[0010] The technical solution of the present application realizes the detection and prediction of slow faults by real-time analysis of performance data; finds the best server node by analyzing healthy server nodes to realize rapid migration of tasks; obtains the corresponding node health score by comprehensive analysis of historical data and current load of server nodes, thereby screening out server nodes that may have unhealthy problems, and selects the best server node through the topological model of network links; predicts whether the relevant server nodes will have slow faults by mining the implicit relationship of related variables in time and space sequences, and cooperates with the single-node slow fault detection algorithm to realize a better slow fault detection method. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 A flowchart of the gray fault handling method provided in an embodiment of the present application; Figure 2 A specific flow chart of the gray fault handling method provided in the embodiment of the present application; Figure 3 A structural diagram of a gray fault handling device provided in an embodiment of the present application; Figure 4 The exemplary systems provided for the embodiments of the present application can be used to implement the various embodiments described in the present application. DETAILED DESCRIPTION

[0013] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0014] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0015] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0016] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the gray fault handling method depends, the specific application environment architecture or specific hardware architecture is described here.

[0017] With the rapid development of artificial intelligence (AI) technology, large-scale distributed training clusters have become the standard for training complex deep learning models. Clusters typically consist of hundreds or thousands of GPU servers capable of processing models with billions or even trillions of parameters. However, during large-scale training, server node failures, especially slow failures, have become a key factor affecting training efficiency. Slow failures, where node performance degrades but not completely fails, can cause delays or even failure of the entire training task. Although slow failures (gray failures) are very common in real-world applications, they are often difficult to quickly detect and effectively mitigate. These failures are caused by a variety of reasons, including hardware performance degradation, network congestion, and resource contention.

[0018] Therefore, it is particularly important to develop a system that can detect and handle slow failures quickly, accurately and automatically.

[0019] However, existing fault detection and recovery methods mostly focus on faults that cause operation to stop, and pay less attention to slow faults that cause implicit performance degradation. At the same time, since the system can still operate normally when a slow fault occurs, general fault detection methods are not suitable for slow fault detection.

[0020] This application proposes a slow fault detection method, which is mainly aimed at slow fault detection in distributed training clusters, reduces manual intervention, significantly improves the speed and accuracy of slow fault detection, reduces the need for manual detection, and improves overall training efficiency and system reliability.

[0021] At the same time, existing detection methods for slow faults are mainly aimed at homogeneous cluster devices such as storage, such as SSD storage and HBM. Slow detection of diverse heterogeneous devices such as AI computing is still in the initial stage.

[0022] This application proposes a slow fault detection method for detecting multi-heterogeneous devices such as AI computing. This method can better detect slow fault devices by combining single-node detection and overall detection.

[0023] The embodiment of the present application provides a gray fault handling method, such as Figure 1 As shown, the method is applied to a gray fault processing system, the system includes a server node cluster, and the method includes: Collect performance data of all server nodes; Detect gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

[0024] It is understood that this application aims to propose a deep learning-based system for adaptively detecting and handling slow failures in large-scale distributed training clusters. This system can monitor performance indicators in real time during training, automatically detect and locate slow failure nodes, and take effective recovery measures to minimize the impact on training tasks. The general framework includes data collection, deep learning detection, fault location, fault recovery, cross-framework integration, real-time performance feedback and optimization, and scalable fault handling strategies. These mechanisms work together to form a complete solution to improve the reliability and efficiency of distributed training clusters.

[0025] Here, gray failure refers to a failure in which the performance of the server node degrades but does not completely fail. Gray failure can cause delays or even failure of the entire training task.

[0026] The embodiment of the present application provides a gray fault handling method, such as Figure 1 As shown, the method includes: Step S01, collecting performance data of all server nodes; Here, this application is mainly aimed at slow fault detection of large-scale deep learning training cluster facilities. The data collection module of this application is divided into a single-node performance data collection module and a global performance data aggregation module.

[0027] Step S011, the system includes a central coordinator, and a monitoring unit is set on each server node; Collecting hardware and network performance data of the server node by a monitoring unit according to a first threshold interval, wherein the hardware and network performance data include graphics processor utilization and serial expansion bus transmission rate; sending the performance data of each server node to the central coordinator via remote communication; The central coordinator analyzes the fault detection iteration time and the communication delay between the server nodes.

[0028] Specifically, in the single-node performance data collection module, this application deploys a lightweight agent program on each server node to monitor the node's local hardware and network performance (such as GPU utilization, PCIe transmission rate) in real time. The agent program sends performance indicator data to the central coordinator through shared memory and remote communication. The central coordinator collects performance data of all server nodes and analyzes the training iteration time and inter-node communication delay. Here, the first threshold interval time is 30 seconds.

[0029] Step S02: Detect server nodes with gray faults in the server cluster based on the performance data of all server nodes and the graph neural network model.

[0030] Specifically, for tasks related to server operation, it should be possible to detect which nodes (server nodes) have slow failures. Therefore, as a central coordinator that collects all performance data, the above data is statistically analyzed to detect whether slow failures have occurred in related server nodes.

[0031] Step S021: creating a graph neural network model, wherein the graph neural network model includes edges and server nodes; According to the performance data of all server nodes, a dynamic spatiotemporal feature graph G(V,E t ), where V is the performance data feature vector of each server node i ;E t is the weight of the communication link between server nodes; Get the performance data feature vector X of each server node i at time point t t , X t ={X 1,t ,X 2,t ,...,X n,t}, n is the number of server nodes; get the learning matrix W q 、W k 、X t Transpose ; By formula: , calculate the influence weight A between server nodes t ; Get the hidden state of the current layer l of the graph neural network model , the learning matrix W of the server node v , the edge learning matrix W e , the weight E of the communication link between server nodes t , the first bias term b; By formula: , calculate the hidden state of the next layer of the graph neural network model ; Get the hidden state of layer m of the graph neural network model , void factor d; By formula: , calculate the high-dimensional vector O that integrates spatiotemporal features t ; The grey failure probability of a single server node is determined through a fully connected linear layer, activation function, and a high-dimensional vector that integrates spatiotemporal features.

[0032] Specifically, the Graph Conv Network (GCN) is mainly composed of edges and nodes. Therefore, GCN training and detection can be used in node clusters. This application models the collected server node cluster performance data information as a dynamic spatiotemporal feature graph G(V,E t ), where V represents the real-time feature vector carried by each server node i ; To better represent the data of V at time point t, here we use X t It represents the real-time feature vector carried by each server node i at time point t, namely X t ={X 1,t ,X 2,t ,...,X n,t}, where n represents the number of server nodes; edge E t Represents the weight of the communication link between nodes i and j, which is composed of real-time communication indicators such as delay, bandwidth, and congested packet count.

[0033] When extracting features, this application is divided into spatial features and temporal features: for spatial features, adaptive adjacent matrix and mixed feature convolution are used. The adaptive adjacent matrix dynamically calculates the influence weights between nodes through the attention mechanism to highlight the key communication links. Activation represents the activation function, and LeakyReLu is mainly used here. q and W k is a learnable matrix, and the final result A t Reflects the real-time dependencies between server nodes (e.g., server groups that communicate frequently have higher weights): ; For mixed feature convolution, this application aggregates node features and edge features at the same time and updates the hidden state as shown in the following formula, where Represents the hidden state of the current layer of GCN, W v and W e They represent the learnable matrices of server nodes and edges respectively, and b is the bias term, which increases the model fitting ability and prevents the output value from approaching 0: ; For temporal features, the TCN network is mainly used for multi-scale temporal modeling to capture the degradation patterns of different temporal granularities, namely , where d=1,2,4 is the hidden state H of the current layer l Different sampling intervals make the output O t Incorporates short-term fluctuations (such as sudden load) with long-term trends (such as hardware aging).

[0034] TCN (Temporal Convolutional Network) is a convolutional network based on causal convolution and dilated convolution, specifically used for modeling time series.

[0035] Causality: Ensure that the output at the current moment depends only on the current and previous time steps; Dilation: Expand the receptive field by controlling the sampling interval to model long-term dependencies; Dilation factors d = 1, 2, 4: represent the different dilation rates used in TCN to achieve multi-scale temporal modeling; Each layer of TCN uses a different void factor to model separately: d=1: short-term time dependence; d=2: medium-term time dependence; d=4: long-term time dependence.

[0036] Step S022, obtaining a high-dimensional vector O of the fusion spatiotemporal features t , fully connected linear layer MLP, activation function Sigmoid; By formula: p i =Sigmoid(MLP( t )), calculate the grey failure probability p of a single server node i ; Determining whether the gray failure probability of a single server node is greater than a first preset value; If so, it is determined that a gray fault occurs on the server node; if not, the server node with a gray fault in the server cluster is re-detected using the cluster gray fault detection algorithm.

[0037] It is understandable that, in the end, the above network outputs the failure probability of a single server node by combining the fully connected layer and the activation function Sigmoid. , the formula is as follows: where MLP is a linear layer, and its dimension and number of layers can be set specifically: p i =Sigmoid(MLP(t )); When p i When the threshold is exceeded (the default threshold here is 0.7), it is considered that a slow fault has occurred and the single-node slow fault detection algorithm needs to be started for further verification.

[0038] Among them, O t In order to integrate the high-dimensional vector of the above-mentioned spatiotemporal features, MLP is used as a linear layer to transform O t The high-dimensional vector is converted into a low-dimensional vector of single logits output of the fault probability, and finally according to the activation function Sigmoid, a probability p with a value between 0 and 1 is output. i .

[0039] Here, the first preset value is 0.7.

[0040] Step S023, updating the learning matrix parameters through the binary cross entropy loss function; The learning matrix parameters are updated through the binary cross entropy loss function, including: Get the learning matrix parameter W, input feature J, and the second bias term b 0 , Sigmoid activation function; By formula: y^=Sigmoid(W×J+b 0 ), calculate the probability y^ of each input feature belonging to the positive class; The learning matrix parameters are updated by the probability that each input feature belongs to the positive class.

[0041] Specifically, this application adopts the binary cross entropy loss function to update the learnable parameter W in the algorithm.

[0042] Step S03 , in response to a gray fault occurring in any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm.

[0043] Specifically, slow fault detection is divided into intra-node slow fault detection and overall slow fault detection; intra-node slow fault detection mainly collects and analyzes relevant data of each server node to detect whether there is significant performance degradation. If there is significant performance degradation, the node task will be migrated to other healthy nodes; overall slow fault detection mainly collects global node performance data to compare the performance differences of each node. When a node is identified as significantly lagging behind, it will be marked.

[0044] Here, slow failure detection within a server node typically involves interaction delays between components within a single server node. For example, if the response time for a service to call another service significantly exceeds a fourth threshold, it is considered a slow failure. The fourth threshold can be set based on historical performance data of the server node, ranging from several hundred milliseconds to several seconds. Detection of overall system slow failures: This refers to performance degradation at the distributed system level. This refers not only to a single service call, but also includes the overall performance of all network requests, message queue processing, etc. in the system. For overall slow failure detection, if the system's average transaction processing time increases to the fifth threshold, it indicates a systemic slow failure problem. The fifth threshold ranges from several seconds to tens of seconds.

[0045] Step S031: Calculate the performance data observation value mean μ and performance data observation value variance σ based on the historical performance data of the first server node. 2 ; Obtain the cumulative deviation S of the indicator of the first server node at time t-1 according to the second threshold interval (20 seconds) t-1 , the observed value x at time t t , performance data observation value mean μ, indicator drift compensation term k, where indicator drift compensation term k=5σ 2 ; By formula: S t =max(0,S t-1 +(x t -μ-k)), calculate the cumulative deviation S of the indicator of the first server node at time t t ; In response to the cumulative deviation of the indicator of any one of the first server nodes at time t being greater than the indicator drift compensation item value, recalculating the cumulative deviation of the indicator of the first server node at time t according to a third threshold interval (1 second); In response to the average value of the cumulative deviation of the indicator of the first server node at time t being greater than the indicator drift compensation item value, it is determined that a gray fault occurs in the first server node.

[0046] Specifically, for the detection of slow faults within the node, a lightweight online real-time detection method is required to enable rapid and accurate detection when a fault occurs. Therefore, this application adopts a cumulative sum algorithm to detect slow faults within the node; since the slow faults of AI training mainly focus on GPU computing and network link related indicators, the calculation here is also mainly for GPU computing and network link transmission rate; first, a period of historical data of the normal operation of the first server node is collected, and the relevant mean μ and variance σ are calculated respectively. 2 , and set the indicator drift compensation term k as the normal fluctuation range ratio (here k=5σ 2), in real-time detection, the cumulative deviation S of the first server node at time t is calculated according to the following formula t : S t =max(0,S t-1 +(x t -μ-k)); The sampling interval is 20 seconds, based on the latest observation value x t , calculate the cumulative deviation of the indicator S t , and S t With k=5σ 2 Compare, when any S t If the value of S is greater than the indicator drift compensation item, the sampling interval is reduced to 1 second and the operation sampling is continued for 10 seconds. t When the average value is greater than the indicator drift compensation item value, it is determined that the current node has a slow failure, and the task is migrated to a healthy node; at the same time, in order to ensure that the mean μ and variance σ are 2 Still representative, therefore, this application uses a sliding window method to update the mean μ and variance σ 2 That is, after running for a certain period of time (such as 20 sampling points and always sliding backward), fitting calculations are performed based on the relevant data collected in the records to ensure the validity of the statistical values.

[0047] Step S032: Obtain the historical GPU utilization of the first server node a , the number of historical GPU utilization data points a; By formula: , calculate the mean μ of the performance data observation value of the first server node; Obtaining a mean μ of the performance data observation values ​​of the first server node; By formula: , calculate the variance σ of the performance data observation value of the first server node 2 .

[0048] Specifically, the historical data of the normal operation of the first server node is collected, and the relevant mean μ and variance σ are calculated respectively. 2 ; Variance σ 2 Indicates the degree of dispersion of the data distribution.

[0049] Step S04 : in response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link.

[0050] Step S041, determining whether there is a gray fault server node in the server node group to be migrated; In response to the fact that there is no gray fault server node in the server node group to be migrated, the server node group to be migrated is screened using the server node health scoring mechanism; The server node cluster to be migrated is screened using the server node health scoring mechanism, including: Obtain the weight coefficient α of the historical failure rate of the server node to be migrated, and the weight coefficient β of the current load centrality of the server node to be migrated; Dynamically adjust the weight coefficient of the historical failure rate of the server node to be migrated and the weight coefficient of the current load centrality of the server node to be migrated; Get the historical failure rate F of the server node to be migrated 历史故障率 , the load occupancy rate F of the server node to be migrated 当前负载 ; By formula: Y i =α(1-F 历史故障率 )+β(1-F 当前负载 ), calculate the health score Y of the server node to be migrated i ; Determine the health score of the server node to be migrated.

[0051] Specifically, during task migration, in addition to ensuring that the server node of the task to be migrated is not a slow fault node, it is also necessary to ensure the suitability of the communication link to ensure the operating efficiency of the node communication; therefore, when selecting the node for task migration, attention should be paid to maintaining information such as the quality of the selected node and the topological structure of the relevant operating link.

[0052] Regarding the quality of the maintenance server node, the following formula can be used to dynamically confirm whether the current server node is suitable for this task, where Y i is the health score of node i, ranging from [0, 1]; F 历史故障率 represents the historical failure rate of node i, ranging from [0, 1]; F 当前负载 represents the load occupancy of the current node i. This occupancy takes the highest current load occupancy, such as GPU utilization and memory occupancy, and ranges from [0, 1]. α and β are weight coefficients reflecting the centrality of the historical failure rate and the current load. To ensure that the total weight of the health score is 1, α + β = 1: Y i =α(1-F 历史故障率 )+β(1-F 当前负载 ); Step S042, determining whether the health score of the server node to be migrated is less than a first preset value; If so, the server node to be migrated is deleted from the server node group to be migrated and is marked; if not, the target migration server node is determined through the topology model of the network link.

[0053] Specifically, when the health score of a server node is lower than the first preset value (0.7), the node is not allowed to accept orders and is fed back to the main controller to be marked. After screening the server nodes to be migrated, it is necessary to further determine how to select the best server node, which is done here through the network link topology and communication rate.

[0054] Step S043: Create a network link topology model C=(Z,R), where Z is the set of server nodes to be migrated. is the set of communication links between the server nodes to be migrated, each communication link e∈R; Obtain the latency (e), bandwidth (e), and CNP (e) of each link e, as well as the importance parameter γ (value range is customizable, here 1-10) for controlling link congestion. By formula: , calculate the index weight B(e) of the communication performance of the server node to be migrated; The server node to be migrated with the smallest indicator weight value of communication performance in the server node group to be migrated is selected as the target migration server node.

[0055] Specifically, this application models the network communication of distributed training as a graph C=(Z,R), where Z is a set of server nodes, representing the nodes in the training cluster (GPU, CPU or computer); R is a set of links, representing the communication paths between server nodes (such as InfiniBand links), and for each link e∈R, the indicator weight B(e) represents an indicator of communication performance.

[0056] Based on the above graph, the server node with the least impact on the communication link is selected and added to the model training as the final target migration server node.

[0057] Step S044: Obtain the task priority parameter R, the standard deviation of the historical failure rate of the server node to be migrated , the standard deviation of the current load of the server node to be migrated ; By formula: , calculate the weight coefficient α of the historical failure rate of the server node to be migrated; By formula: , calculate the weight coefficient β of the current load centrality of the server node to be migrated; When the task priority parameter is the second preset value (0), the historical failure rate of the server node to be migrated is ignored; When the task priority parameter is the third preset value (1), the current load condition of the server node to be migrated is ignored.

[0058] Specifically, how to determine α and β should be dynamically adjusted according to the actual scenario. For example, in a high-load system, the weight of the current load should be given priority, while in tasks with high reliability requirements, the weight of the historical failure rate should be given priority. Therefore, in order to adapt to the dynamic needs of different scenarios, this application adjusts the above two weights through a dynamic method: for different tasks, a task priority parameter R can be set in the range of [0,1] to indicate the degree of reliability requirements of the task. When R=0, the historical failure rate is completely ignored; when R=1, the current load is completely ignored. Therefore, the specific formula is as follows, where is the standard deviation of the historical failure rate, is the standard deviation of the current load, where the historical failure rate is calculated based on the probability of all failures that occurred before the current server node, and the current load is calculated based on the mean of the current task across all nodes where it is deployed: ; ; The above method combines the global state and task requirements to provide a more flexible weight adjustment method to better screen and determine relevant nodes.

[0059] Step S05: Migrate the computing task of the first server node to the target migration server node.

[0060] Step S051, deleting the first server node from the communication link of the task calculation; When the computing task of the first server node is assigned to the target migration server node, the migration task of the server node without gray fault is interrupted; Optimize network communication link resources and verify the consistency of network link topology model parameters after migration.

[0061] Specifically, task migration aims to quickly migrate affected tasks from gray-failed nodes to healthy nodes while minimizing the impact of migration on overall system performance. After detecting a slow failure, it is usually necessary to remove the slow node from the critical path of task computation to prevent the anomaly from affecting the normal operation of other tasks. Simultaneously, migrated tasks need to be reallocated to the best-performing server nodes to optimize the utilization of computing and network resources, while ensuring the consistency of model parameters and the correctness of the training process after migration.

[0062] At the same time, to avoid full restart technologies such as checkpoints, only the computing tasks on slow nodes are migrated during migration. During the migration process, the remaining healthy nodes of this task will suspend training until the migration task is completed.

[0063] like Figure 2 As shown, the technical solution of the present application collects the performance data of all server nodes; detects the server nodes with gray faults in the server cluster based on the performance data of all server nodes and the graph neural network model; in response to a gray fault in any server node in the server cluster, any server is set as the first server node, and the gray fault of the first server node is verified through a single-node gray fault detection algorithm; in response to the gray fault verification of the first server node being correct, the target migration server node is determined through the server node health scoring mechanism and the topological model of the network link; and the computing task of the first server node is migrated to the target migration server node. The present application significantly improves the speed and accuracy of gray fault detection in the server cluster, reduces the need for manual detection, and improves the reliability of the overall system.

[0064] In addition, network communication link resources are optimized and the consistency of the parameters of the migrated network link topology model is verified, including: Optimize network communication link resources, including: Use MPLS-TE methods to dynamically adjust traffic paths to avoid congestion and improve bandwidth utilization; Implementing SDN (Software Defined Network)-based solutions to flexibly manage data plane routing through a centralized control plane; Distribute traffic across multiple links to reduce the risk of single point overload and use Anycast software to provide users with the nearest service node; Set priorities based on application requirements to ensure sufficient bandwidth and service quality for critical businesses; Establish redundant paths and quickly switch to backup links when the primary link fails, maintaining service continuity; For large-scale data center networks, adopt energy-saving mode to shut down some unused links or devices during low-load periods; Verify the consistency of the migrated network link topology model parameters, including: Deploy automated test scripts to compare network configuration files before and after migration to check for any differences. Collect status information of network devices through SNMP protocol and monitor changes in performance indicators; In distributed systems, especially in areas such as caching and load balancing, consistent hashing algorithms are used to maintain the consistency of data distribution and reduce the amount of redistributed data when the network topology changes.

[0065] By combining the above technologies and methods, network communication link resources can be effectively optimized and the new network link topology model parameters can be ensured to be consistent with expectations after network migration, thereby ensuring network stability and efficiency.

[0066] The gray fault handling method provided in the embodiment of the present application can be further improved and optimized without departing from the technical solution of the present application, and these improvements and optimizations should also be considered as the scope of protection of the present application.

[0067] The beneficial effects of the technical solution provided by the embodiments of the present application are: This application significantly improves the speed and accuracy of gray fault detection in server clusters, reduces the need for manual detection, and improves the reliability of the overall system.

[0068] The technical solution of the present application realizes the detection and prediction of slow faults by real-time analysis of performance data; finds the best server node by analyzing healthy server nodes to realize rapid migration of tasks; obtains the corresponding node health score by comprehensive analysis of historical data and current load of server nodes, thereby screening out server nodes that may have unhealthy problems, and selects the best server node through the topological model of network links; predicts whether the relevant server nodes will have slow faults by mining the implicit relationship of related variables in time and space sequences, and cooperates with the single-node slow fault detection algorithm to realize a better slow fault detection method.

[0069] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0070] The embodiment of the present application also provides a gray fault processing device, such as Figure 3 As shown, the device includes: a collection module, a first detection module, a second detection module, a determination module, and a processing module.

[0071] In this embodiment, the collection module is used to collect performance data of all server nodes; The first detection module is used to detect server nodes with gray faults in the server cluster based on the performance data of all server nodes and the graph neural network model; a second detection module configured to, in response to a gray fault occurring on any server node in the server cluster, set any server as a first server node and verify the gray fault of the first server node using a single-node gray fault detection algorithm; a determination module configured to determine a target migration server node based on a server node health scoring mechanism and a network link topology model in response to a gray fault check of the first server node being correct; A processing module is used to migrate the computing task of the first server node to the target migration server node.

[0072] In this embodiment, the acquisition module is used to set a monitoring unit on each server node; Collecting hardware and network performance data of the server node by the monitoring unit according to a first threshold interval, wherein the hardware and network performance data include graphics processor utilization and serial expansion bus transmission rate; Sending performance data of each server node to the central coordinator via remote communication; The fault detection iteration time and the communication delay between server nodes are analyzed through the central coordinator.

[0073] In one embodiment, the first detection module is configured to create a graph neural network model, wherein the graph neural network model includes edges and server nodes; According to the performance data of all server nodes, a dynamic spatiotemporal feature graph G(V,E t ), where V is the performance data feature vector of each server node i , E t is the weight of the communication link between server nodes; Get the performance data feature vector X of each server node i at time point t t , X t ={X 1,t ,X 2,t ,...,X n,t}, n is the number of server nodes; get the learning matrix W q 、W k 、X t Transpose ; By formula: , calculate the influence weight A between server nodes t ; Get the hidden state of the current layer l of the graph neural network model , the learning matrix W of the server node v , the edge learning matrix W e, the weight E of the communication link between server nodes t , the first bias term b; By formula: , calculate the hidden state of the next layer of the graph neural network model ; Get the hidden state of layer m of the graph neural network model , void factor d; By formula: , calculate the high-dimensional vector O that integrates spatiotemporal features t ; The grey failure probability of a single server node is determined through a fully connected linear layer, activation function, and a high-dimensional vector that integrates spatiotemporal features.

[0074] In one embodiment, the first detection module is used to obtain a high-dimensional vector O that integrates spatiotemporal features. t , fully connected linear layer MLP, activation function Sigmoid; By formula: p i =Sigmoid(MLP( t )), calculate the grey failure probability p of a single server node i ; Determining whether the gray failure probability of a single server node is greater than a first preset value; If so, it is determined that a gray fault occurs on the server node; if not, the server node with a gray fault in the server cluster is re-detected using the cluster gray fault detection algorithm.

[0075] In one embodiment, the second detection module is used to calculate the performance data observation value mean μ and the performance data observation value variance σ based on the historical performance data of the first server node. 2 ; Obtain the cumulative deviation S of the indicator of the first server node at time t-1 according to the second threshold interval t-1 , the observed value x at time t t , performance data observation value mean μ, indicator drift compensation term k, where indicator drift compensation term k=5σ 2 ; By formula: S t =max(0,S t-1 +(x t -μ-k)), calculate the cumulative deviation S of the indicator of the first server node at time t t ; In response to the cumulative deviation of the indicator of any one of the first server nodes at time t being greater than the indicator drift compensation item value, recalculating the cumulative deviation of the indicator of the first server node at time t according to a third threshold interval; In response to the average value of the cumulative deviation of the indicator of the first server node at time t being greater than the indicator drift compensation item value, it is determined that a gray fault occurs in the first server node.

[0076] In one embodiment, the second detection module is used to obtain the historical graphics processor utilization gpu of the first server node. a , the number of historical GPU utilization data points a; By formula: , calculate the mean μ of the performance data observation value of the first server node; Obtaining a mean μ of the performance data observation values ​​of the first server node; By formula: , calculate the variance σ of the performance data observation value of the first server node 2 .

[0077] In one embodiment, the determination module is used to determine whether there is a server node with a gray fault in the server node group to be task migrated; In response to the fact that there is no gray fault server node in the server node group to be migrated, the server node group to be migrated is screened using the server node health scoring mechanism; The server node cluster to be migrated is screened using the server node health scoring mechanism, including: Obtain the weight coefficient α of the historical failure rate of the server node to be migrated, and the weight coefficient β of the current load centrality of the server node to be migrated; Dynamically adjust the weight coefficient of the historical failure rate of the server node to be migrated and the weight coefficient of the current load centrality of the server node to be migrated; Get the historical failure rate F of the server node to be migrated 历史故障率 , the load occupancy rate F of the server node to be migrated 当前负载 ; By formula: Y i =α(1-F 历史故障率 )+β(1-F 当前负载 ), calculate the health score Y of the server node to be migrated i ; Determine the health score of the server node to be migrated.

[0078] In one embodiment, the determination module is configured to determine whether the health score of the server node to be task migrated is less than a first preset value; If so, the server node to be migrated is deleted from the server node group to be migrated and is marked; if not, the target migration server node is determined through the topology model of the network link.

[0079] In one embodiment, the determination module is used to create a network link topology model C=(Z,R), where Z is a set of server node clusters to be migrated, R is a set of communication links between the server nodes to be migrated, and each communication link e∈R; Obtain the latency (e) of each link e, the bandwidth (e) of each link e, the congestion notification packet count (CNP) of each link e, and the importance parameter γ that controls link congestion. By formula: , calculate the index weight B(e) of the communication performance of the server node to be migrated; The server node to be migrated with the smallest indicator weight value of communication performance in the server node group to be migrated is selected as the target migration node.

[0080] In one embodiment, the processing module is configured to remove the first server node from the communication link for task computation; When the computing task of the first server node is assigned to the target migration server node, the migration task of the server node without gray fault is interrupted; Optimize network communication link resources and verify the consistency of network link topology model parameters after migration.

[0081] In one embodiment, the first detection module is used to update the learning matrix parameters through a binary cross entropy loss function; The learning matrix parameters are updated through the binary cross entropy loss function, including: Get the learning matrix parameter W, input feature J, and the second bias term b 0 , Sigmoid function; By formula: y^=Sigmoid(W×J+b 0 ), calculate the probability y^ of each input feature belonging to the positive class; The learning matrix parameters are updated by the probability that each input feature belongs to the positive class.

[0082] In one embodiment, a determination module is used to obtain a task priority parameter , the standard deviation of the historical failure rate of the server node to be migrated , the standard deviation of the current load of the server node to be migrated ; By formula: , calculate the weight coefficient α of the historical failure rate of the server node to be migrated; By formula: , calculate the weight coefficient β of the current load centrality of the server node to be migrated; When the task priority parameter is the second preset value, the historical failure rate of the server node to be migrated is ignored; When the task priority parameter is the third preset value, the current load condition of the server node to be migrated is ignored.

[0083] The beneficial effects of the technical solution provided by the embodiments of the present application are: This application significantly improves the speed and accuracy of gray fault detection in server clusters, reduces the need for manual detection, and improves the reliability of the overall system.

[0084] The technical solution of the present application realizes the detection and prediction of slow faults by real-time analysis of performance data; finds the best server node by analyzing healthy server nodes to realize rapid migration of tasks; obtains the corresponding node health score by comprehensive analysis of historical data and current load of server nodes, thereby screening out server nodes that may have unhealthy problems, and selects the best server node through the topological model of network links; predicts whether the relevant server nodes will have slow faults by mining the implicit relationship of related variables in time and space sequences, and cooperates with the single-node slow fault detection algorithm to realize a better slow fault detection method.

[0085] For the description of the features in the embodiment corresponding to the gray fault handling device, please refer to the relevant description of the embodiment corresponding to the gray fault handling method, which will not be repeated here.

[0086] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps of an embodiment of a gray fault processing method, the method comprising: Collect performance data of all server nodes; Detect gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

[0087] like Figure 4 As shown, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the steps in the embodiment of the gray fault handling method when running, the method comprising: Collect performance data of all server nodes; Detect gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

[0088] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0089] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the gray fault processing method embodiment are implemented. The method includes: Collect performance data of all server nodes; Detect gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

[0090] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps in the embodiment of the gray fault processing method, the method including: Collect performance data of all server nodes; Detect gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, any server is set as a first server node, and the gray fault of the first server node is verified using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

[0091] This application significantly improves the speed and accuracy of gray fault detection in server clusters, reduces the need for manual detection, and improves the reliability of the overall system.

[0092] The technical solution of the present application realizes the detection and prediction of slow faults by real-time analysis of performance data; finds the best server node by analyzing healthy server nodes to realize rapid migration of tasks; obtains the corresponding node health score by comprehensive analysis of historical data and current load of server nodes, thereby screening out server nodes that may have unhealthy problems, and selects the best server node through the topological model of network links; predicts whether the relevant server nodes will have slow faults by mining the implicit relationship of related variables in time and space sequences, and cooperates with the single-node slow fault detection algorithm to realize a better slow fault detection method.

[0093] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0094] The above is a detailed introduction to the gray fault handling method, device, medium and program product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.

Claims

1. A gray fault handling method, characterized in that: The method is applied to a gray fault processing system, the system including a server node cluster, and the method includes: Collect performance data of all server nodes; Detecting gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, setting the any server as a first server node, and verifying the gray fault of the first server node using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrate the computing task of the first server node to the target migration server node.

2. The gray fault processing method according to claim 1, characterized in that: The system includes a central coordinator, and the collection of performance data of all server nodes includes: Set up a monitoring unit on each server node; collecting hardware and network performance data of the server node by the monitoring unit according to a first threshold interval, wherein the hardware and network performance data include graphics processor utilization and serial expansion bus transmission rate; sending the performance data of each server node to the central coordinator via remote communication; The central coordinator analyzes the fault detection iteration time and the communication delay between the server nodes.

3. The gray fault processing method according to claim 1, characterized in that: The detecting of gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model includes: Creating a graph neural network model, wherein the graph neural network model includes edges and server nodes; According to the performance data of all server nodes, a dynamic spatiotemporal feature graph G(V,E t ), where V is the performance data feature vector of each server node i ;E t is the weight of the communication link between server nodes; Get the performance data feature vector X of each server node i at time point t t , X t ={X 1,t ,X 2,t ,...,X n,t }, n is the number of server nodes; get the learning matrix W q 、W k 、X t Transpose ; By formula: , calculate the influence weight A between server nodes t ; Get the hidden state of the current layer l of the graph neural network model , the learning matrix W of the server node v , the edge learning matrix W e , the weight E of the communication link between server nodes t , the first bias term b; By formula: , calculate the hidden state of the next layer of the graph neural network model ; Get the hidden state of layer m of the graph neural network model , void factor d; By formula: , calculate the high-dimensional vector O that integrates spatiotemporal features t ; The grey failure probability of a single server node is determined through a fully connected linear layer, activation function, and a high-dimensional vector that integrates spatiotemporal features.

4. The gray fault processing method according to claim 3, characterized in that: The gray failure probability of a single server node is determined by using a fully connected linear layer, an activation function, and a high-dimensional vector that integrates spatiotemporal features, including: Get the high-dimensional vector O that integrates spatiotemporal features t , fully connected linear layer MLP, activation function Sigmoid; By formula: p i =Sigmoid(MLP( t )), calculate the grey failure probability p of a single server node i ; Determining whether the gray failure probability of the single server node is greater than a first preset value; If so, it is determined that a gray fault occurs on the server node; if not, the server node that has a gray fault in the server cluster is re-detected using a cluster gray fault detection algorithm.

5. The gray fault processing method according to claim 1, characterized in that: The verifying the gray fault of the first server node by using a single-node gray fault detection algorithm includes: Calculate the performance data observation value mean μ and the performance data observation value variance σ according to the historical performance data of the first server node 2 ; Obtain the cumulative deviation S of the indicator of the first server node at time t-1 according to the second threshold interval t-1 , the observed value x at time t t , performance data observation value mean μ, indicator drift compensation term k, where the indicator drift compensation term k=5σ 2 ; By formula: S t =max(0,S t-1 +(x t -μ-k)), calculate the cumulative deviation S of the indicator of the first server node at time t t ; In response to the cumulative deviation of the indicator of any one of the first server nodes at time t being greater than the indicator drift compensation item value, recalculating the cumulative deviation of the indicator of the first server node at time t according to a third threshold interval; In response to the average value of the cumulative deviation of the indicator of the first server node at time t being greater than the indicator drift compensation item value, it is determined that a gray fault occurs in the first server node.

6. The gray fault processing method according to claim 5, characterized in that: The calculating, based on the historical performance data of the first server node, a mean of the performance data observation value and a variance of the performance data observation value, includes: Get the historical GPU utilization of the first server node a , the number of historical GPU utilization data points a; By formula: , calculating a mean μ of the performance data observation values ​​of the first server node; Obtaining a mean μ of the performance data observation values ​​of the first server node; By formula: , calculate the variance σ of the performance data observation value of the first server node 2 .

7. The gray fault processing method according to claim 1, characterized in that: The determining of the target migration server node by using the server node health scoring mechanism and the network link topology model includes: Determine whether there are any gray faulty server nodes in the server node group to be migrated; In response to the absence of a gray fault server node in the server node group to be migrated, screening the server node group to be migrated by using a server node health scoring mechanism; The screening of the server node group to be migrated using the server node health scoring mechanism includes: Obtain the weight coefficient α of the historical failure rate of the server node to be migrated, and the weight coefficient β of the current load centrality of the server node to be migrated; Dynamically adjust the weight coefficient of the historical failure rate of the server node to be migrated and the weight coefficient of the current load centrality of the server node to be migrated; Get the historical failure rate F of the server node to be migrated 历史故障率 , the load occupancy rate F of the server node to be migrated 当前负载 ; By formula: Y i =α(1-F 历史故障率 )+β(1-F 当前负载 ), calculate the health score Y of the server node to be migrated i ; The health score of the server node to be migrated is determined.

8. The gray fault processing method according to claim 7, characterized in that: The determining of the health score of the server node to be migrated includes: Determine whether the health score of the server node to be migrated is less than a first preset value; If so, the server node to be migrated is deleted from the server node group to be migrated, and the server node to be migrated is marked; if not, the target migration server node is determined through the topology model of the network link.

9. The gray fault processing method according to claim 8, characterized in that: The determining of the target migration server node by using the topology model of the network link includes: Create a network link topology model C = (Z, R), where Z is the set of server nodes to be migrated, R is the set of communication links between server nodes to be migrated, and each communication link e∈R; Obtain the latency (e) of each link e, the bandwidth (e) of each link e, the congestion notification packet count (CNP) of each link e, and the importance parameter γ that controls link congestion. By formula: , calculate the index weight B(e) of the communication performance of the server node to be migrated; The server node to be migrated with the smallest indicator weight value of communication performance in the group of server nodes to be migrated is used as the target migration node.

10. The gray fault processing method according to claim 1, characterized in that: Migrating the computing task of the first server node to the target migration server node includes: Deleting the first server node from the communication link of task computing; When the computing task of the first server node is assigned to the target migration server node, the migration task of the server node without gray fault is interrupted; Optimize network communication link resources and verify the consistency of network link topology model parameters after migration.

11. The gray fault processing method according to claim 3, characterized in that: The method comprises: The learning matrix parameters are updated through the binary cross entropy loss function; The updating of the learning matrix parameters by the binary cross entropy loss function includes: Get the learning matrix parameter W, input feature J, and the second bias term b 0 , Sigmoid function; By formula: y^=Sigmoid(W×J+b 0 ), calculate the probability y^ of each input feature belonging to the positive class; The learning matrix parameters are updated by the probability that each input feature belongs to the positive class.

12. The gray fault processing method according to claim 7, characterized in that: The dynamically adjusting the weight coefficient of the historical failure rate of the server node to be migrated and the weight coefficient of the current load centrality of the server node to be migrated includes: Get the task priority parameter R and the standard deviation of the historical failure rate of the server node to be migrated , the standard deviation of the current load of the server node to be migrated ; By formula: , calculate the weight coefficient α of the historical failure rate of the server node to be migrated; By formula: , calculate the weight coefficient β of the current load centrality of the server node to be migrated; When the task priority parameter is the second preset value, the historical failure rate of the server node to be migrated is ignored; When the task priority parameter is the third preset value, the current load condition of the server node to be migrated is ignored.

13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the gray fault handling method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the gray fault processing method according to any one of claims 1 to 12.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the gray fault processing method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Micro-service system fault positioning method based on graph neural network

    CN114721860A

  • Fault detection method and device, electronic equipment and storage medium

    CN117573459A

  • Cross-server fault prediction system and method based on transfer learning

    CN117950965A

  • Method and device for realizing power grid hidden fault diagnosis processing based on graph convolutional network, processor and computer readable storage medium thereof

    CN120009663A

  • Timely software defect prediction method and system based on deep learning

    US20250225057A1

Cited By

  • Fault processing method and system for modular mainboard expansion interface, electronic equipment and storage medium

    CN121301067A