A gray failure handling method, device, medium, and program product

Through the graph neural network model and health scoring mechanism, combined with the network link topology model, computing tasks are automatically detected and migrated, solving the problem of detecting and handling gray faults in distributed training clusters, achieving fast and accurate fault detection and task migration, and improving system reliability.

CN120631741BActive Publication Date: 2025-10-17INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511129420.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-17
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing technologies have difficulty in quickly and accurately detecting and handling slow failures in distributed training clusters, especially gray failures of multiple heterogeneous devices, which lead to training task delays or failures, and rely on manual detection, which is inefficient.

Method used

A graph neural network model is combined with a single-node gray fault detection algorithm. By monitoring the performance data of server nodes in real time, the health scoring mechanism and network link topology model are used to automatically detect and migrate computing tasks, reducing manual intervention.

Benefits of technology

It significantly improves the speed and accuracy of gray fault detection, reduces the need for manual inspection, and improves system reliability and overall training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631741B_ABST
    Figure CN120631741B_ABST
Patent Text Reader

Abstract

The application discloses a gray fault processing method and device, medium and program product, relates to the technical field of detection, and comprises the following steps: collecting performance data of all server nodes; detecting server nodes with gray faults in a server cluster according to the performance data of all server nodes and a graph neural network model; in response to the occurrence of a gray fault in any one server node in the server cluster, verifying the gray fault of the first server node through a single-node gray fault detection algorithm; in response to the correct verification of the gray fault of the first server node, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; and migrating the computing task of the first server node to the target migration server node. The application significantly improves the gray fault detection speed and accuracy, reduces the need for manual detection, and improves the reliability of the overall system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of detection, and in particular to a gray fault processing method, device, medium and program product. BACKGROUND

[0002] At present, the related fault processing method mainly relies on manual or automatic detection of hardware redundancy after fault modeling, but these methods are mainly designed for faults that cause the system to stop, and have not good detection effect for slow faults (gray faults) that implicitly reduce system performance; for example, hardware redundancy may cause "gray faults", that is, the performance of the system gradually decreases unknowingly, and manual detection requires a large amount of human resources and is difficult to respond in real time. SUMMARY

[0003] The present application provides a gray fault processing method, device, medium and program product, the method comprising collecting performance data of all server nodes; detecting server nodes with gray faults in a server cluster according to the performance data of all server nodes and a graph neural network model; in response to any one server node in the server cluster having a gray fault, setting the any one server as a first server node, and verifying the gray fault of the first server node through a single-node gray fault detection algorithm; in response to the gray fault verification of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of network links; and migrating the computing task of the first server node to the target migration server node. The present application significantly improves the gray fault detection speed and accuracy, reduces the need for manual detection, and improves the reliability of the overall system.

[0004] The present application provides a gray fault processing method, the method being applied to a gray fault processing system, the system comprising a server node cluster, and the method comprising:

[0005] collecting performance data of all server nodes;

[0006] detecting server nodes with gray faults in a server cluster according to the performance data of all server nodes and a graph neural network model;

[0007] in response to any one server node in the server cluster having a gray fault, setting the any one server as a first server node, and verifying the gray fault of the first server node through a single-node gray fault detection algorithm;

[0008] in response to the gray fault verification of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of network links.

[0009] migrating the computing task of the first server node to the target migration server node.

[0010] The application further provides a gray fault processing device, comprising:

[0011] a collecting module configured to collect performance data of all server nodes;

[0012] a first detecting module configured to detect a server node with a gray fault in a server cluster according to the performance data of all server nodes and a graph neural network model;

[0013] a second detecting module configured to, in response to any one server node in the server cluster having a gray fault, set the any one server as a first server node and check the gray fault of the first server node by using a single-node gray fault detection algorithm;

[0014] a determining module configured to, in response to the gray fault checking of the first server node being correct, determine a target migration server node by using a server node health score mechanism and a topology model of a network link;

[0015] a processing module configured to migrate the computing task of the first server node to the target migration server node.

[0016] The application further provides an electronic device, comprising a memory configured to store a computer program and a processor configured to execute the computer program to implement the steps of the gray fault processing method, wherein the method comprises:

[0017] collecting performance data of all server nodes;

[0018] detecting a server node with a gray fault in a server cluster according to the performance data of all server nodes and a graph neural network model;

[0019] in response to any one server node in the server cluster having a gray fault, setting the any one server as a first server node and checking the gray fault of the first server node by using a single-node gray fault detection algorithm;

[0020] in response to the gray fault checking of the first server node being correct, determining a target migration server node by using a server node health score mechanism and a topology model of a network link;

[0021] migrating the computing task of the first server node to the target migration server node.

[0022] The application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program.

[0023] Collect performance data of all server nodes;

[0024] Detect a server node with a gray fault in the server cluster according to the performance data of all server nodes and the graph neural network model;

[0025] In response to the occurrence of a gray fault in any one of the server nodes in the server cluster, set the any one server as a first server node, and verify the gray fault of the first server node by using a single-node gray fault detection algorithm;

[0026] In response to the correct verification of the gray fault of the first server node, determine a target migration server node by using a server node health score mechanism and a topology model of a network link;

[0027] Migrate the computing task of the first server node to the target migration server node.

[0028] The application further provides a computer program product, and the computer program product comprises a computer program.

[0029] Collect performance data of all server nodes;

[0030] Detect a server node with a gray fault in the server cluster according to the performance data of all server nodes and the graph neural network model;

[0031] In response to the occurrence of a gray fault in any one of the server nodes in the server cluster, set the any one server as a first server node, and verify the gray fault of the first server node by using a single-node gray fault detection algorithm;

[0032] In response to the correct verification of the gray fault of the first server node, determine a target migration server node by using a server node health score mechanism and a topology model of a network link;

[0033] Migrate the computing task of the first server node to the target migration server node.

[0034] Through the present application, since the method comprises collecting performance data of all server nodes; detecting a server node with a gray fault in the server cluster according to the performance data of all server nodes and a graph neural network model; in response to any one server node in the server cluster having a gray fault, setting the any one server as a first server node, and verifying the gray fault of the first server node through a single-node gray fault detection algorithm; in response to the gray fault verification of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; and migrating the computing task of the first server node to the target migration server node. The present application significantly improves the gray fault detection speed and accuracy, reduces the need for manual detection, and improves the reliability of the overall system.

[0035] The technical scheme of the present application realizes detection and prediction of slow faults through real-time analysis of performance data; finds out the best server node through analysis of healthy server nodes to realize fast migration of tasks; obtains the corresponding node health score through comprehensive analysis of historical data and current load of the server nodes, so as to screen out server nodes that may have health problems, and selects the best server node through a topology model of a network link; predicts whether the related server nodes will have slow faults through mining of the implicit relationship of related variables in time and space sequences, and cooperates with the single-node slow fault detection algorithm to realize a better slow fault detection method. BRIEF DESCRIPTION OF DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0037] Figure 1 A gray fault processing method flowchart provided for the embodiments of the present application;

[0038] Figure 2 A gray fault processing method specific flowchart provided for the embodiments of the present application;

[0039] Figure 3 A structure diagram of a gray fault processing device provided for the embodiments of the present application;

[0040] Figure 4 An exemplary system that can be used to implement the various embodiments described in the present application is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0041] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0042] It should be noted that in the description of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0043] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0044] In combination with the specific application environment architecture or the specific hardware architecture on which the execution of the gray fault processing method depends, the specific application environment architecture or the specific hardware architecture is described here.

[0045] With the rapid development of artificial intelligence technology, large-scale distributed training clusters have become the standard for training complex deep learning models; clusters are usually composed of hundreds or thousands of graphics processing unit (GPU) servers, which can handle models with tens of billions or even thousands of billions of parameters; however, in the process of large-scale training, the problem of server node failure, especially slow failure, has become a key factor affecting training efficiency. Slow failure, i.e. the failure of node performance decline but not complete failure, will cause delay and even failure of the entire training task; although slow failure (gray failure) is very common in practical applications, it is often difficult to be quickly detected and effectively mitigated, and these failures are caused by various reasons, including hardware performance degradation, network congestion, resource contention, etc.

[0046] Therefore, it is particularly important to develop a system that can quickly, accurately and automatically detect and handle slow failures.

[0047] However, existing fault detection and recovery mostly target at failures that cause running to stop, and pay less attention to slow failures that cause implicit performance decline. At the same time, since slow failures occur when the system is still running normally, general fault detection methods are not suitable for slow failure detection.

[0048] The application proposes a slow fault detection method, mainly aiming at slow fault detection in a distributed training cluster, reducing manual intervention, significantly improving slow fault detection speed and accuracy, reducing the need for manual detection, and improving overall training efficiency and system reliability.

[0049] At the same time, the existing detection method for slow faults mainly targets storage of such homogeneous cluster devices, such as SSD storage, HBM, etc. The slow detection of such multi-element heterogeneous devices as AI computing is still in its initial stage.

[0050] The application proposes a slow fault detection method for detecting AI computing such multi-element heterogeneous devices. This method combines single-node detection and overall detection to better detect slow fault devices.

[0051] Embodiments of the application provide a gray fault processing method, as shown in Figure 1 The method is applied to a gray fault processing system, which includes a server node cluster. The method includes:

[0052] Collecting performance data of all server nodes;

[0053] Detecting server nodes with gray faults in the server cluster according to the performance data of all server nodes and a graph neural network model;

[0054] In response to any one server node in the server cluster having a gray fault, the any one server is set as a first server node, and a single-node gray fault detection algorithm is used to verify the gray fault of the first server node;

[0055] In response to the gray fault verification of the first server node being correct, a target migration server node is determined through a server node health scoring mechanism and a topology model of a network link;

[0056] The computing task of the first server node is migrated to the target migration server node.

[0057] It can be understood that the application aims to propose a large-scale distributed training cluster based on deep learning for slow fault adaptive detection and processing system. The system can monitor performance indicators in real time during the training process, automatically detect and locate slow fault nodes, and take effective recovery measures to minimize the impact on the training task. The general framework includes data collection, deep learning detection, fault location, fault recovery, cross-framework integration, real-time performance feedback and optimization, and scalable fault handling strategies. These mechanisms work together to form a complete solution to improve the reliability and efficiency of distributed training clusters.

[0058] Here, the gray fault refers to a fault in which the performance of the server node decreases but is not completely disabled, and the gray fault causes delay or even failure of the entire training task.

[0059] Embodiments of the present application provide a gray fault processing method, as shown in the method comprises: Figure 1

[0060] Step S01, collecting performance data of all server nodes;

[0061] Here, the present application mainly aims at slow fault detection of large-scale deep learning training cluster facilities, and the data collection module of the present application is divided into a single node performance data collection module and a global performance data summary module.

[0062] Step S011, the system comprises a central coordinator, and a monitoring unit is arranged on each server node;

[0063] The monitoring unit collects hardware and network performance data of the server node according to a first threshold interval time, wherein the hardware and network performance data includes GPU utilization rate and PCIe transmission rate;

[0064] The performance data of each server node is sent to the central coordinator through remote communication;

[0065] The central coordinator analyzes the communication delay between the fault detection iteration time and the server node.

[0066] Specifically, in the single node performance data collection module, the present application deploys a lightweight agent program on each server node to monitor the local hardware and network performance (such as GPU utilization rate and PCIe transmission rate) in real time. The agent program sends performance index data to the central coordinator through shared memory and remote communication, and the central coordinator collects performance data of all server nodes, analyzes the training iteration time and the inter-node communication delay, and the first threshold interval time is 30 seconds.

[0067] Step S02, detecting the server node in which the gray fault occurs in the server cluster according to the performance data of all server nodes and the graph neural network model.

[0068] Specifically, for related tasks running on the server, it should be detected as much as possible which nodes (server nodes) have slow faults, therefore, as the central coordinator that collects all performance data, statistical analysis is performed on the above data to detect whether the related server nodes have slow faults.

[0069] Step S021, creating a graph neural network model, wherein the graph neural network model comprises edges and server nodes; ​

[0070] According to the performance data of all server nodes, a dynamic space-time feature graph G(V, E t ) is established, wherein V is a performance data feature vector of each server node i ; E t is a weight of a communication link between server nodes;

[0071] A performance data feature vector X t of each server node i at a time point t is obtained, X t ={X 1,t , X 2,t ,..., X n,t}, n is the number of server nodes; a learning matrix W q , W k , and a transpose of X t are obtained; ;

[0072] An influence weight A t between server nodes is calculated through a formula: ;

[0073] A hidden state H of a current layer l of a graph neural network model is obtained, a learning matrix W v of a server node, a learning matrix W e of an edge, a weight E t of a communication link between server nodes, and a first bias term b are obtained;

[0074] A hidden state H of a next layer of the graph neural network model is calculated through a formula: ;

[0075] A hidden state H of an m layer of the graph neural network model is obtained, and a cavity factor d is obtained;

[0076] A high-dimensional vector O t fused with space-time features is calculated through a formula: ;

[0077] A gray failure probability of a single server node is determined through a fully connected linear layer, an activation function, and the high-dimensional vector fused with space-time features.

[0078] Specifically, a graph convolution network (GCN) is mainly composed of edges and nodes, and therefore, the GCN can be used for training detection in a node cluster; the performance data information of the collected server node cluster is modeled as a dynamic space-time feature graph G(V, E t ), wherein V represents a real-time feature vector carried by each server node i ; To better represent the data of V at time point t, here X t represents the real-time feature vector carried by each server node i at time point t, i.e. X t ={X 1,t ,X 2,t ,...,X n,t}, where n represents the number of server nodes; the edge E t represents the weight of the communication link between i and j nodes, which is composed of real-time communication indicators such as delay, bandwidth, and congestion packet count.

[0079] In feature extraction, the present application is divided into spatial features and temporal features: for spatial features, adaptive adjacency matrix and hybrid feature convolution are used, the adaptive adjacency matrix dynamically calculates the weight between nodes through attention mechanism, highlighting key communication links, where Activation represents the activation function, here mainly using LeakyReLu, W q and W k are learnable matrices, and the final result A t reflects the real-time dependence between server nodes (such as server groups with frequent communication have higher weights):

[0080] ;

[0081] For hybrid feature convolution, the present application simultaneously aggregates node features and edge features to update hidden states, as shown in the following formula, where represents the hidden state of the current layer of GCN, W v and W e are learnable matrices for server nodes and edges respectively, and b is the bias term, which increases the model fitting ability and avoids the output value approaching 0:

[0082] ;

[0083] For temporal features, TCN network is mainly used for multi-scale time modeling to capture different time granularity degradation patterns, i.e. , where d = 1, 2, 4 is the different sampling interval of the current layer hidden state H l , so that the output O t fuses short-term fluctuations (such as sudden load) and long-term trends (such as hardware aging).

[0084] TCN (Temporal Convolutional Network) is a convolutional network based on causal convolution (Causal Convolution) and dilated convolution (Dilated Convolution), which is specifically used for modeling time series.

[0085] Causality: Ensure that the output at the current moment depends only on the current and previous time steps;

[0086] Dilation: Expand the receptive field by controlling the sampling interval to model long-term dependencies;

[0087] Dilation factors d = 1, 2, 4: represent the different dilation rates used in TCN to achieve multi-scale temporal modeling;

[0088] Each layer of TCN uses a different void factor to model separately:

[0089] d=1: short-term time dependence; d=2: medium-term time dependence; d=4: long-term time dependence.

[0090] Step S022, obtaining a high-dimensional vector O of the fusion spatiotemporal features t , fully connected linear layer MLP, activation function Sigmoid;

[0091] By formula: p i =Sigmoid(MLP( t )), calculate the grey failure probability p of a single server node i ;

[0092] Determining whether the gray failure probability of a single server node is greater than a first preset value;

[0093] If so, it is determined that a gray fault occurs on the server node; if not, the server node with a gray fault in the server cluster is re-detected using the cluster gray fault detection algorithm.

[0094] It is understandable that, in the end, the above network outputs the failure probability of a single server node by combining the fully connected layer and the activation function Sigmoid. , the formula is as follows: where MLP is a linear layer, and its dimension and number of layers can be set specifically:

[0095] p i =Sigmoid(MLP( t ));

[0096] When p i When the threshold is exceeded (the default threshold here is 0.7), it is considered that a slow fault has occurred and the single-node slow fault detection algorithm needs to be started for further verification.

[0097] Among them, O t In order to integrate the high-dimensional vector of the above-mentioned spatiotemporal features, MLP is used as a linear layer to transform O tThe high-dimensional vector is converted into a low-dimensional vector of single logits output of failure possibility, and finally a value p in 0-1 is output according to the activation function Sigmoid i .

[0098] Here, the first preset value is 0.7.

[0099] In step S023, the learning matrix parameter is updated by a binary cross-entropy loss function.

[0100] The learning matrix parameter is updated by a binary cross-entropy loss function, including:

[0101] Obtain the learning matrix parameter W, the input feature J, and the second bias item b 0 , and the Sigmoid activation function.

[0102] Calculate the probability y^ of each input feature belonging to the positive class by the formula: y^ = Sigmoid (W x J + b 0 ).

[0103] Update the learning matrix parameter according to the probability of each input feature belonging to the positive class.

[0104] Specifically, the learning parameter W in the binary cross-entropy loss function update algorithm is adopted.

[0105] In step S03, in response to the occurrence of gray failure of any server node in the server cluster, any server is set as the first server node, and the gray failure of the first server node is verified by a single-node gray failure detection algorithm.

[0106] Specifically, the slow failure detection is divided into node internal slow failure detection and overall slow failure detection; the node internal slow failure detection mainly collects and analyzes relevant data of each server node to detect whether there is a significant performance decline, and if there is a significant performance decline, the node task is migrated to other healthy nodes; the overall slow failure detection mainly collects global node performance data to compare the performance differences of each node, and when a node is identified to be significantly behind, it is labeled.

[0107] Here, for the slow failure detection within the server node: it usually involves the interaction delay between the internal components of a single server node; for example, if the response time of a service calling another service significantly exceeds the fourth threshold value, it is considered to be a slow failure; the fourth threshold value can be set according to the historical performance data of the server node, wherein the fourth threshold value is several hundred milliseconds to several seconds.

[0108] For detection of slow failure of the whole system: refers to the performance decline of the whole distributed system level, not limited to single service call, but including the overall performance of all network requests, message queue processing, etc. For the detection of overall slow failure, if the average transaction processing time of the system increases to the fifth threshold value, it indicates that there is a systemic slow failure problem, wherein the fifth threshold value is several seconds to tens of seconds.

[0109] Step S031, calculating the performance data observation value mean μ and the performance data observation value variance σ according to the historical performance data of the first server node 2 ;

[0110] According to the second threshold interval time (20 seconds), the index cumulative deviation S of the first server node at time t-1 is obtained t-1 , the observation value x t of time t, the performance data observation value mean μ, and the index drift compensation term k, wherein the index drift compensation term k=5σ 2 ;

[0111] According to the formula: S t =max(0,S t-1 +(x t -μ-k)), the index cumulative deviation S t of the first server node at time t is calculated.

[0112] In response to any one of the index cumulative deviations of the first server node at time t being greater than the index drift compensation term value, the index cumulative deviation of the first server node at time t is recalculated according to the third threshold interval time (1 second);

[0113] In response to the average value of the index cumulative deviation of the first server node at time t being greater than the index drift compensation term value, it is determined that the first server node has a gray failure.

[0114] Specifically, for the slow failure detection of the node, a lightweight online real-time detection method is needed to quickly and accurately detect the failure when it occurs. Therefore, the cumulative sum algorithm is used to detect the slow failure of the node. Since the slow failure of AI training mainly focuses on GPU operation and network link related indicators, this calculation is mainly for GPU operation and network link transmission rate. First, collect a period of historical data of the first server node running normally, calculate the related mean μ and variance σ 2 , and set the index drift compensation term k as the normal fluctuation range proportion (here k=5σ 2 ). In real-time detection, the index cumulative deviation S t of the first server node at time t is calculated according to the following formula:

[0115] St = max(0, S t-1 + (x t - μ - k));

[0116] The sampling interval is 20 seconds, and the index cumulative deviation S t is calculated according to the current latest observation value x t , and S t is compared with k=5σ 2 , when any one S t is greater than the index drift compensation term value, the sampling interval is reduced to 1 second, and the operation is continued for 10 seconds, when the average value of S t in this period of time is greater than the index drift compensation term value, it is determined that the current node has a slow fault, and the task is migrated to a healthy node; At the same time, in order to ensure that the average value μ and the variance σ 2 are still representative during a large load time, therefore, the application adopts the sliding window method to update the average value μ and the variance σ 2 , that is, after a certain period of time (such as 20 sampling points, and always sliding backward), the relevant data recorded are fitted and calculated to ensure the effectiveness of the statistical value.

[0117] Step S032, obtaining the historical graphics processor utilization rate gpu a of the first server node, the number of historical graphics processor utilization rate data points a;

[0118] The performance data observation value mean μ of the first server node is calculated by the formula:

[0119] The performance data observation value mean μ of the first server node is obtained;

[0120] The performance data observation value variance σ 2 of the first server node is calculated by the formula:

[0121] Specifically, the historical data of the normal operation of the first server node is collected, and the relevant mean μ and variance σ 2 are calculated respectively; the variance σ 2 represents the dispersion degree of data distribution.

[0122] Step S04, in response to the correct gray fault check of the first server node, the target migration server node is determined through the server node health score mechanism and the topology model of the network link.

[0123] Step S041, judging whether there is a server node with gray fault in the server node group to be migrated; ​​

[0124] In response to no gray fault server node appearing in the server node group to be migrated, the server node group to be migrated is screened through a health score mechanism of the server node;

[0125] Screening the server node group to be migrated through the health score mechanism of the server node includes:

[0126] Obtaining a weight coefficient α of a historical failure rate of the server node to be migrated and a weight coefficient β of a current load centrality of the server node to be migrated;

[0127] Dynamically adjusting the weight coefficient of the historical failure rate of the server node to be migrated and the weight coefficient of the current load centrality of the server node to be migrated;

[0128] Obtaining a historical failure rate F 历史故障率 of the server node to be migrated and a load occupancy rate F 当前负载 of the server node to be migrated;

[0129] Calculating a health score Y i of the server node to be migrated through a formula: Y i =α(1-F 历史故障率 )+β(1-F 当前负载 );

[0130] Judging the health score of the server node to be migrated.

[0131] Specifically, in task migration, in addition to ensuring that the server node to be migrated is not a slow fault node, the suitability of the communication link also needs to be ensured to ensure the running efficiency of node communication; therefore, on the selected node of task migration, attention should be paid to maintaining the good and bad of the selected node and the topology structure of the related running link and the like information.

[0132] For maintaining the good and bad of the server node, the current server node can be dynamically confirmed as suitable for the task through the following formula, wherein Y i is the health score of node i, ranging from 0 to 1; F 历史故障率 represents the historical failure rate of node i, ranging from 0 to 1; F 当前负载 represents the load occupancy rate of the current node i, and the occupancy rate takes the highest current load occupancy, such as GPU utilization rate, memory occupancy rate and the like, ranging from 0 to 1; α and β are weight coefficients reflecting the historical failure rate and the centrality of the current load, and in order to ensure that the total weight of the health score is 1, therefore α+β=1:

[0133] Y i =α(1-F 历史故障率 )+β(1-F 当前负载 );

[0134] Step S042, judge whether the health score of the to-be-task-migrated server node is less than the first preset value;

[0135] If yes, the to-be-task-migrated server node is deleted from the to-be-task-migrated server node group, and the to-be-task-migrated server node is marked; if no, the target migration server node is determined through the topology model of the network link.

[0136] Specifically, when the health score of the server node is lower than the first preset value (0.7), the node is not allowed to accept orders and is fed back to the total controller for marking; after the to-be-task-migrated server node is screened, it is necessary to further determine how to select the best server node, which is determined by the network link topology and communication rate.

[0137] Step S043, create a topology model C=(Z,R) of the network link, Z is a set of to-be-task-migrated server node groups, is a set of communication links between to-be-task-migrated server nodes, and each communication link e∈R;

[0138] Obtain the delay Latency(e) of each link e, the bandwidth Bandwidth(e) of each link e, CNP(e), and the importance parameter γ (the value range is defined and set, and here is 1-10) of controlling link congestion;

[0139] Calculate the index weight B(e) of the communication performance of the to-be-task-migrated server node through the formula:

[0140] The to-be-task-migrated server node with the minimum index weight value of the communication performance in the to-be-task-migrated server node group is selected as the target migration server node.

[0141] Specifically, the network communication of distributed training is modeled as a graph C=(Z,R), where Z is a set of server nodes, representing nodes (GPUs, CPUs or computers) in the training cluster; R is a set of links, representing communication paths (such as InfiniBand links) between server nodes, each link e∈R, and the index weight B(e) represents the index of communication performance.

[0142] According to the above graph, the server node with the minimum communication link influence is selected as the final target migration server node to join the model training.

[0143] Step S044, obtain the task priority parameter R, the standard deviation of the historical failure rate of the to-be-task-migrated server node , and the standard deviation of the current load of the to-be-task-migrated server node ;​

[0144] The weight coefficient a of the historical failure rate of the task migration server node to be migrated is calculated by the formula:

[0145] The weight coefficient β of the current load centrality of the task migration server node to be migrated is calculated by the formula:

[0146] When the task priority parameter is the second preset value (0), the historical failure rate of the task migration server node to be migrated is ignored.

[0147] When the task priority parameter is the third preset value (1), the current load of the task migration server node to be migrated is ignored.

[0148] Specifically, how to determine a and β should be dynamically adjusted according to the actual scene. For example, in a high-load system, the weight of the current load should be given priority, and in a task with high reliability requirements, the weight of the historical failure rate should be given priority. Therefore, in order to adapt to the dynamic needs of different scenes, the application adjusts the above two weights by a dynamic method: for different tasks, a task priority parameter R can be set, ranging from [0, 1], to represent the degree of requirement for reliability, when R=0, the historical failure rate is completely ignored; R=1 completely ignores the current load, therefore, the specific formula is as follows, wherein is the standard deviation of the historical failure rate, is the standard deviation of the current load, wherein the historical failure rate is calculated according to the probability of all previous failures of the current server node, and the current load is calculated according to the mean value of all nodes to which the current task is deployed:

[0149] ;

[0150] ;

[0151] The above method combines the global state and task demand to provide a more flexible weight adjustment method to better select and determine the relevant node.

[0152] Step S05, migrating the computing task of the first server node to the target migration server node.

[0153] Step S051, deleting the first server node from the communication link of the task computing;

[0154] When the computing task of the first server node is assigned to the target migration server node, the migration task of the gray failure server node is interrupted;

[0155] ​​Optimize network communication link resources, and check the consistency of the network link topology model parameters after migration.

[0156] Specifically, task migration aims to quickly migrate the affected tasks from the node with gray failure to a healthy node while minimizing the impact of migration on the overall performance of the system. After detecting a slow failure, it is usually necessary to remove the slow node from the critical path of task computation to avoid abnormal impact on the normal operation of other tasks. At the same time, the migrated tasks need to be redistributed to the best server node to optimize the utilization of computing and network resources and ensure the consistency of the model parameters after migration and the correctness of the training process.

[0157] At the same time, in order to avoid techniques such as checkpointing that require a full restart, only the computing tasks on the slow node are migrated during migration. During the migration process, the remaining healthy nodes of the task will suspend training until the migration task is completed.

[0158] As shown in Figure 2 The technical scheme of the present application collects performance data of all server nodes; detects the server node with gray failure in the server cluster according to the performance data of all server nodes and the graph neural network model; in response to the occurrence of gray failure in any server node in the server cluster, sets any server node as a first server node, and checks the gray failure of the first server node through a single-node gray failure detection algorithm; in response to the correct gray failure check of the first server node, determines a target migration server node through a server node health scoring mechanism and a topology model of network link; and migrates the computing tasks of the first server node to the target migration server node. The gray failure detection of the present application in the server cluster significantly improves the speed and accuracy of gray failure detection, reduces the need for manual detection, and improves the reliability of the overall system.

[0159] In addition, the network communication link resources are optimized, and the consistency of the network link topology model parameters after migration is checked, including:

[0160] The network communication link resources are optimized, including:

[0161] The MPLS-TE method is used to dynamically adjust the traffic path to avoid congestion and improve bandwidth utilization;

[0162] Implement an SDN (Software Defined Network) based solution to flexibly manage the routing of the data plane through a centralized control plane;

[0163] Distribute traffic over multiple links to reduce the risk of single-point overload and use Anycast software to provide the nearest service node for users;

[0164] Set priorities according to application requirements to ensure that critical services have sufficient bandwidth and quality of service guarantees;

[0165] Establish redundant paths and quickly switch to backup links when the primary link fails to maintain service continuity;

[0166] For large-scale data center networks, use energy-saving mode to close part of the unused links or devices during low load period;

[0167] Check the consistency of the network link topology model parameters after migration, including:

[0168] Deploy automated test scripts to compare network configuration files before and after migration and check for differences;

[0169] Collect network device status information through SNMP protocol and monitor changes in performance indicators;

[0170] In distributed systems, especially in caching, load balancing and other fields, use consistent hashing algorithm to maintain data distribution consistency and reduce the amount of data redistribution when network topology changes.

[0171] By combining the above techniques and methods, network communication link resources can be effectively optimized, and the new network link topology model parameters after network migration can be ensured to be consistent with expectations, thereby ensuring the stability and efficiency of the network.

[0172] The gray fault handling method provided by the embodiments of the present application can be improved and optimized without departing from the technical solutions of the present application, and these improvements and optimizations should also be considered within the protection scope of the present application.

[0173] The technical solutions provided by the embodiments of the present application have the following beneficial effects:

[0174] The present application significantly improves the speed and accuracy of gray fault detection in server clusters, reduces the need for manual detection, and improves the reliability of the overall system.

[0175] The technical solutions of the present application realize slow fault detection and prediction by real-time analysis of performance data; find the best server node through analysis of healthy server nodes to realize fast migration of tasks; obtain the corresponding node health score through comprehensive analysis of historical data and current load of server nodes, so as to screen out server nodes that may have health problems, and select the best server node through the topology model of network link; predict whether related server nodes will have slow faults by mining the implicit relationship of related variables in time and space sequences, and cooperate with single-node slow fault detection algorithm to realize better slow fault detection method.

[0176] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software on a general hardware platform as necessary, and of course can also be realized by hardware, but in many cases the former is a better embodiment.

[0177] The embodiments of the present application also provide a gray fault processing device, as shown in the figure, the device comprises a collection module, a first detection module, a second detection module, a determination module and a processing module. Figure 3

[0178] In the embodiment, the collection module is configured to collect performance data of all server nodes.

[0179] The first detection module is configured to detect a server node with a gray fault in a server cluster according to the performance data of all server nodes and a graph neural network model.

[0180] The second detection module is configured to, in response to the occurrence of a gray fault in any one of the server nodes in the server cluster, set the any one server node as a first server node, and check the gray fault of the first server node by using a single-node gray fault detection algorithm.

[0181] The determination module is configured to, in response to the correct checking of the gray fault of the first server node, determine a target migration server node by using a server node health score mechanism and a topology model of a network link.

[0182] The processing module is configured to migrate a computing task of the first server node to the target migration server node.

[0183] In the embodiment, the collection module is configured to set a monitoring unit on each server node.

[0184] The monitoring unit is configured to collect hardware and network performance data of the server node according to a first threshold interval time, wherein the hardware and network performance data include a graphics processor utilization rate and a serial advanced technology attachment transmission rate.

[0185] The performance data of each server node is sent to a central coordinator by remote communication.

[0186] The central coordinator is configured to analyze a fault detection iteration time and a communication delay between the server nodes.

[0187] In one of the embodiments, the first detection module is configured to create a graph neural network model, wherein the graph neural network model comprises edges and server nodes.

[0188] ​According to the performance data of all server nodes, a dynamic space-time feature map G(V, E t ) is established, wherein V is the performance data feature vector of each server node i , and E t is the weight of the communication link between the server nodes.

[0189] The performance data feature vector X t of each server node i at time point t is obtained, X t ={X 1,t ,X 2,t ,...,X n,t}, and n is the number of server nodes; the learning matrix W q , the transpose of W k , and X t are obtained.

[0190] The influence weight A t between the server nodes is calculated by the formula: .

[0191] The hidden state H of the current layer l of the graph neural network model is obtained, the learning matrix W v of the server nodes, the learning matrix W e of the edges, the weight E t of the communication link between the server nodes, and the first bias term b.

[0192] The hidden state H of the next layer of the graph neural network model is calculated by the formula: .

[0193] The hidden state H of the m-th layer of the graph neural network model is obtained, and the empty factor d.

[0194] The high-dimensional vector O t fused with the space-time features is calculated by the formula: .

[0195] The gray failure probability of a single server node is determined by a fully connected linear layer, an activation function, and the high-dimensional vector fused with the space-time features.

[0196] In one embodiment, the first detection module is configured to obtain the high-dimensional vector O t fused with the space-time features, a fully connected linear layer MLP, and an activation function Sigmoid.

[0197] The gray failure probability p of a single server node is calculated by the formula: i p=Sigmoid(MLP(O t )).​i ;

[0198] Determining whether the gray failure probability of a single server node is greater than a first preset value;

[0199] If so, it is determined that a gray fault occurs on the server node; if not, the server node with a gray fault in the server cluster is re-detected using the cluster gray fault detection algorithm.

[0200] In one embodiment, the second detection module is used to calculate the performance data observation value mean μ and the performance data observation value variance σ based on the historical performance data of the first server node. 2 ;

[0201] Obtain the cumulative deviation S of the indicator of the first server node at time t-1 according to the second threshold interval t-1 , the observed value x at time t t , performance data observation value mean μ, indicator drift compensation term k, where indicator drift compensation term k=5σ 2 ;

[0202] By formula: S t =max(0,S t-1 +(x t -μ-k)), calculate the cumulative deviation S of the indicator of the first server node at time t t ;

[0203] In response to the cumulative deviation of the indicator of any one of the first server nodes at time t being greater than the indicator drift compensation item value, recalculating the cumulative deviation of the indicator of the first server node at time t according to a third threshold interval;

[0204] In response to the average value of the cumulative deviation of the indicator of the first server node at time t being greater than the indicator drift compensation item value, it is determined that a gray fault occurs in the first server node.

[0205] In one embodiment, the second detection module is used to obtain the historical graphics processor utilization gpu of the first server node. a , the number of historical GPU utilization data points a;

[0206] By formula: , calculate the mean μ of the performance data observation value of the first server node;

[0207] Obtaining a mean μ of the performance data observation values ​​of the first server node;

[0208] By formula: , calculate the variance σ of the performance data observation value of the first server node 2 .

[0209] In one embodiment, the determining module is configured to determine whether there is a gray fault server node in the server node group to be migrated;

[0210] In response to that there is no gray fault server node in the server node group to be migrated, the server node group to be migrated is screened through a health score mechanism of the server node;

[0211] The screening of the server node group to be migrated through the health score mechanism of the server node comprises:

[0212] obtaining a weight coefficient α of a historical failure rate of the server node to be migrated and a weight coefficient β of a current load centrality of the server node to be migrated;

[0213] dynamically adjusting the weight coefficient of the historical failure rate of the server node to be migrated and the weight coefficient of the current load centrality of the server node to be migrated;

[0214] obtaining a historical failure rate F 历史故障率 of the server node to be migrated and a load occupancy rate F 当前负载 of the server node to be migrated;

[0215] calculating a health score Y i of the server node to be migrated through a formula: Y i =α(1-F 历史故障率 )+β(1-F 当前负载 );

[0216] judging the health score of the server node to be migrated.

[0217] In one embodiment, the determining module is configured to determine whether the health score of the server node to be migrated is less than a first preset value;

[0218] If yes, the server node to be migrated is deleted from the server node group to be migrated, and the server node to be migrated is marked; if no, a target migration server node is determined through a topology model of a network link.

[0219] In one embodiment, the determining module is configured to create a topology model C=(Z,R) of the network link, Z is a set of the server node group to be migrated, R is a set of communication links between the server nodes to be migrated, and each communication link e∈R;

[0220] obtaining a delay Latency(e) of each link e, a bandwidth Bandwidth(e) of each link e, a congestion notification packet count CNP(e) of each link e, and an importance parameter γ of controlling link congestion.

[0221] The index weight B(e) of the communication performance of the to-be-task-migrated server node is calculated by the formula:

[0222] The to-be-task-migrated server node with the minimum index weight value of the communication performance in the to-be-task-migrated server node group is taken as the target migration node.

[0223] In one of the embodiments, the processing module is configured to delete the first server node from the communication link of the task computing;

[0224] When the computing task of the first server node is allocated to the target migration server node, the migration task of the gray fault server node is interrupted;

[0225] The network communication link resources are optimized, and the consistency of the network link topology model parameters after migration is verified.

[0226] In one of the embodiments, the first detection module is configured to update the learning matrix parameters by using a binary cross-entropy loss function;

[0227] The updating of the learning matrix parameters by using the binary cross-entropy loss function comprises:

[0228] obtaining the learning matrix parameters W, the input features J, and the second bias item b 0 , a Sigmoid function;

[0229] The probability y^ of each input feature belonging to a positive class is calculated by the formula: y^ = Sigmoid(W x J + b 0 );

[0230] The learning matrix parameters are updated by using the probability of each input feature belonging to a positive class.

[0231] In one of the embodiments, the determination module is configured to obtain the task priority parameters , the standard deviation of the historical failure rate of the to-be-task-migrated server node , the standard deviation of the current load of the to-be-task-migrated server node ;

[0232] The weight coefficient a of the historical failure rate of the to-be-task-migrated server node is calculated by the formula:

[0233] The weight coefficient b of the current load centrality of the to-be-task-migrated server node is calculated by the formula:

[0234] ​​​When the task priority parameter is the second preset value, the historical failure rate of the server node to which the task is to be migrated is ignored.

[0235] When the task priority parameter is the third preset value, the current load condition of the server node to which the task is to be migrated is ignored.

[0236] The technical scheme provided by the embodiment of the application has the following beneficial effects:

[0237] The application significantly improves the speed and accuracy of gray failure detection in the server cluster, reduces the need for manual detection, and improves the reliability of the overall system.

[0238] The technical scheme of the application realizes the detection and prediction of slow failure by analyzing performance data in real time; finds out the best server node through analysis of healthy server nodes to realize fast migration of tasks; obtains the corresponding node health score through comprehensive analysis of historical data and current load of the server node, so as to screen out server nodes that may have health problems, and selects the best server node through the topology model of the network link; predicts whether the related server node will have a slow failure by mining the implicit relationship of related variables in time and space sequences, and cooperates with the single-node slow failure detection algorithm to realize a better slow failure detection method.

[0239] The description of the features in the embodiment corresponding to the gray failure processing device can be referred to the related description of the embodiment corresponding to the gray failure processing method, which will not be repeated here.

[0240] The embodiment of the application also provides an electronic device, including a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in the gray failure processing method embodiment, the method includes:

[0241] Collecting performance data of all server nodes;

[0242] Detecting server nodes having gray failure in the server cluster according to the performance data of all server nodes and the graph neural network model;

[0243] In response to the occurrence of gray failure in any one of the server nodes in the server cluster, setting the any one of the server nodes as a first server node, and verifying the gray failure of the first server node through a single-node gray failure detection algorithm;

[0244] In response to the correct verification of the gray failure of the first server node, determining a target migration server node through a server node health score mechanism and a topology model of a network link;

[0245] migrate the computing task of the first server node to the target migration server node.

[0246] As shown in Figure 4 The embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores a computer program, wherein the computer program is configured to execute the steps in the gray fault processing method embodiment when running, and the method comprises the following steps:

[0247] collecting performance data of all server nodes;

[0248] detecting a server node with a gray fault in a server cluster according to the performance data of all server nodes and a graph neural network model;

[0249] In response to the occurrence of a gray fault in any one of the server nodes in the server cluster, setting the any one server as a first server node, and verifying the gray fault of the first server node through a single-node gray fault detection algorithm;

[0250] In response to the correct verification of the gray fault of the first server node, determining a target migration server node through a server node health score mechanism and a topology model of a network link;

[0251] migrating the computing task of the first server node to the target migration server node.

[0252] In an example embodiment, the computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0253] The embodiments of the present application also provide a computer program product, and the computer program product comprises a computer program, and the computer program is executed by a processor to implement the steps in the gray fault processing method embodiment, and the method comprises the following steps:

[0254] collecting performance data of all server nodes;

[0255] detecting a server node with a gray fault in a server cluster according to the performance data of all server nodes and a graph neural network model;

[0256] In response to the occurrence of a gray fault in any one of the server nodes in the server cluster, setting the any one server as a first server node, and verifying the gray fault of the first server node through a single-node gray fault detection algorithm;

[0257] In response to the gray fault of the first server node being verified to be correct, a target migration server node is determined through a server node health scoring mechanism and a topology model of network links.

[0258] The computing task of the first server node is migrated to the target migration server node.

[0259] Embodiments of the present application also provide another computer program product, comprising a non-volatile computer readable storage medium, the non-volatile computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps in the gray fault processing method embodiment, the method comprising:

[0260] Performance data of all server nodes is collected.

[0261] The server nodes in the server cluster that have a gray fault are detected according to the performance data of all server nodes and a graph neural network model.

[0262] In response to any one of the server nodes in the server cluster having a gray fault, the any one server node is set as a first server node, and a single-node gray fault detection algorithm is used to verify the gray fault of the first server node.

[0263] In response to the gray fault of the first server node being verified to be correct, a target migration server node is determined through a server node health scoring mechanism and a topology model of network links.

[0264] The computing task of the first server node is migrated to the target migration server node.

[0265] The present application significantly improves the speed and accuracy of gray fault detection in the server cluster, reduces the need for manual detection, and improves the reliability of the overall system.

[0266] The present application detects and predicts slow faults by analyzing performance data in real time; finds the best server node through analysis of healthy server nodes to achieve fast migration of tasks; obtains the node health score through comprehensive analysis of historical data and current load of the server node, thereby screening out server nodes that may have health problems, and selecting the best server node through a topology model of network links; predicts whether related server nodes will have slow faults by mining the implicit relationship of related variables in time and space sequences, and cooperates with the single-node slow fault detection algorithm to achieve a better slow fault detection method.

[0267] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0268] The above is a detailed introduction to the gray fault handling method, device, medium and program product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.

Claims

1. A gray fault handling method, characterized in that: The method is applied to a gray fault processing system, the system including a server node cluster, and the method includes: Collect performance data of all server nodes; Detecting gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model; In response to a gray fault occurring on any server node in the server cluster, setting the any server as a first server node, and verifying the gray fault of the first server node using a single-node gray fault detection algorithm; In response to the gray fault check of the first server node being correct, determining a target migration server node through a server node health scoring mechanism and a topology model of a network link; Migrating the computing task of the first server node to the target migration server node; The detecting of gray faulty server nodes in the server cluster based on the performance data of all server nodes and the graph neural network model includes: Creating a graph neural network model, wherein the graph neural network model includes edges and server nodes; According to the performance data of all server nodes, a dynamic spatiotemporal feature graph G(V,E t ), where V is the performance data feature vector of each server node i ;E t is the weight of the communication link between server nodes; Get the performance data feature vector X of each server node i at time point t t , X t ={X 1,t ,X 2,t ,...,X n,t }, n is the number of server nodes; get the learning matrix W q 、W k 、X t Transpose ; By formula: , calculate the influence weight A between server nodes t ; Get the hidden state of the current layer l of the graph neural network model , the learning matrix W of the server node v , the edge learning matrix W e , the weight E of the communication link between server nodes t , the first bias term b; By formula: , calculate the hidden state of the next layer of the graph neural network model ; Get the hidden state of layer m of the graph neural network model , void factor d; By formula: , calculate the high-dimensional vector O that integrates spatiotemporal features t ; The grey failure probability of a single server node is determined through a fully connected linear layer, activation function, and a high-dimensional vector that integrates spatiotemporal features.

2. The gray fault processing method according to claim 1, characterized in that: The system includes a central coordinator, and the collection of performance data of all server nodes includes: Set up a monitoring unit on each server node; collecting hardware and network performance data of the server node by the monitoring unit according to a first threshold interval, wherein the hardware and network performance data include graphics processor utilization and serial expansion bus transmission rate; sending the performance data of each server node to the central coordinator via remote communication; The central coordinator analyzes the fault detection iteration time and the communication delay between the server nodes.

3. The gray fault processing method according to claim 1, characterized in that: The gray failure probability of a single server node is determined by using a fully connected linear layer, an activation function, and a high-dimensional vector that integrates spatiotemporal features, including: Get the high-dimensional vector O that integrates spatiotemporal features t , fully connected linear layer MLP, activation function Sigmoid; By formula: p i =Sigmoid(MLP( t )), calculate the grey failure probability p of a single server node i ; Determining whether the gray failure probability of the single server node is greater than a first preset value; If so, it is determined that a gray fault occurs on the server node; if not, the server node that has a gray fault in the server cluster is re-detected using a cluster gray fault detection algorithm.

4. The gray fault processing method according to claim 1, characterized in that: The verifying the gray fault of the first server node by using a single-node gray fault detection algorithm includes: Calculate the performance data observation value mean μ and the performance data observation value variance σ according to the historical performance data of the first server node 2 ; Obtain the cumulative deviation S of the indicator of the first server node at time t-1 according to the second threshold interval t-1 , the observed value x at time t t , performance data observation value mean μ, indicator drift compensation term k, where the indicator drift compensation term k=5σ 2 ; By formula: S t =max(0,S t-1 +(x t -μ-k)), calculate the cumulative deviation S of the indicator of the first server node at time t t ; In response to the cumulative deviation of the indicator of any one of the first server nodes at time t being greater than the indicator drift compensation item value, recalculating the cumulative deviation of the indicator of the first server node at time t according to a third threshold interval; In response to the average value of the cumulative deviation of the indicator of the first server node at time t being greater than the indicator drift compensation item value, it is determined that a gray fault occurs in the first server node.

5. The gray fault processing method according to claim 4, characterized in that: The calculating, based on the historical performance data of the first server node, a mean of the performance data observation value and a variance of the performance data observation value, includes: Get the historical GPU utilization of the first server node a , the number of historical GPU utilization data points a; By formula: , calculating a mean μ of the performance data observation values ​​of the first server node; Obtaining a mean μ of performance data observation values ​​of the first server node; By formula: , calculate the variance σ of the performance data observation value of the first server node 2 .

6. The gray fault processing method according to claim 1, characterized in that: The determining of the target migration server node by using the server node health scoring mechanism and the network link topology model includes: Determine whether there are any gray faulty server nodes in the server node group to be migrated; In response to the absence of a gray fault server node in the server node group to be migrated, screening the server node group to be migrated by using a server node health scoring mechanism; The screening of the server node group to be migrated using the server node health scoring mechanism includes: Obtain the weight coefficient α of the historical failure rate of the server node to be migrated, and the weight coefficient β of the current load centrality of the server node to be migrated; Dynamically adjust the weight coefficient of the historical failure rate of the server node to be migrated and the weight coefficient of the current load centrality of the server node to be migrated; Get the historical failure rate F of the server node to be migrated 历史故障率 , the load occupancy rate F of the server node to be migrated 当前负载 ; By formula: Y i =α(1-F 历史故障率 )+β(1-F 当前负载 ), calculate the health score Y of the server node to be migrated i ; The health score of the server node to be migrated is determined.

7. The gray fault processing method according to claim 6, characterized in that: The determining of the health score of the server node to be migrated includes: Determine whether the health score of the server node to be migrated is less than a first preset value; If so, the server node to be migrated is deleted from the server node group to be migrated, and the server node to be migrated is marked; if not, the target migration server node is determined through the topology model of the network link.

8. The gray fault processing method according to claim 7, characterized in that: The determining of the target migration server node by using the topology model of the network link includes: Create a network link topology model C = (Z, R), where Z is the set of server nodes to be migrated, R is the set of communication links between server nodes to be migrated, and each communication link e∈R; Obtain the latency (e) of each link e, the bandwidth (e) of each link e, the congestion notification packet count (CNP) of each link e, and the importance parameter γ that controls link congestion. By formula: , calculate the index weight B(e) of the communication performance of the server node to be migrated; The server node to be migrated with the smallest indicator weight value of communication performance in the group of server nodes to be migrated is used as the target migration node.

9. The gray fault processing method according to claim 1, characterized in that: Migrating the computing task of the first server node to the target migration server node includes: Deleting the first server node from the communication link of task calculation; When the computing task of the first server node is assigned to the target migration server node, the migration task of the server node without gray fault is interrupted; Optimize network communication link resources and verify the consistency of network link topology model parameters after migration.

10. The gray fault processing method according to claim 1, characterized in that: The method comprises: The learning matrix parameters are updated through the binary cross entropy loss function; The updating of the learning matrix parameters by the binary cross entropy loss function includes: Get the learning matrix parameter W, input feature J, and the second bias term b 0 , Sigmoid function; By formula: y^=Sigmoid(W×J+b 0 ), calculate the probability y^ of each input feature belonging to the positive class; The learning matrix parameters are updated by the probability that each input feature belongs to the positive class.

11. The gray fault processing method according to claim 6, characterized in that: The dynamically adjusting the weight coefficient of the historical failure rate of the server node to be migrated and the weight coefficient of the current load centrality of the server node to be migrated includes: Get the task priority parameter R and the standard deviation of the historical failure rate of the server node to be migrated , the standard deviation of the current load of the server node to be migrated ; By formula: , calculate the weight coefficient α of the historical failure rate of the server node to be migrated; By formula: , calculate the weight coefficient β of the current load centrality of the server node to be migrated; When the task priority parameter is the second preset value, the historical failure rate of the server node to be migrated is ignored; When the task priority parameter is the third preset value, the current load condition of the server node to be migrated is ignored.

12. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the gray fault handling method according to any one of claims 1 to 11 when executing the computer program.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the gray fault processing method according to any one of claims 1 to 11 are implemented.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the gray fault processing method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Micro-service system fault positioning method based on graph neural network

    CN114721860A

  • Fault detection method and device, electronic equipment and storage medium

    CN117573459A