A method and device for active defense of faults, electronic equipment and storage medium

CN122802332APending Publication Date: 2026-09-22SUIYUAN INTELLIGENT TECH (CHENGDU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611135513.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]但是现有技术中在进行集群故障处理时,通常仅能感知硬件层的故障信号,而对可能造成硬件故障发生的软件错误信号无法进行捕捉,因此感知维度缺失,错过了故障提前干预的最佳时间窗口;而在进行故障处理时,仅是从单个计算节点独立分析处理,忽略了节点间故障传播效应,从而降低了故障识别的准确度;在识别出故障后也仅仅是围绕资源预留和任务重调度,防御策略单一

Benefits of technology

[0009]本发明的技术方案,通过构建更新动态故障传播图谱,将硬件的软错误信号与通信拓扑深度融合,并根据故障在动态故障传播图谱的传播路径,预测节点集群中各计算节点的风险等级,并采用与风险等级所匹配的防御策略,进行差异化故障防御,从而准确高效的对故障提前进行干预,保证分布式集群任务的正常执行。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802332A_ABST
    Figure CN122802332A_ABST
Patent Text Reader

Abstract

This invention discloses a proactive fault defense method, apparatus, device, and storage medium, comprising: updating a dynamic fault propagation map by combining a link status signal reported by a network sensing agent with a soft error signal reported by a node agent when a soft error signal is received; predicting the risk level of each computing node based on the updated dynamic fault propagation map; determining a defense strategy matching the risk level for each computing node; and executing the defense strategy to defend against faults in the computing node. By constructing and updating the dynamic fault propagation map, the soft error signal of the hardware is deeply integrated with the communication topology, and the risk level of each computing node in the node cluster is predicted based on the propagation path of the fault in the dynamic fault propagation map. Differentiated fault defense is then performed by adopting a defense strategy matching the risk level, thereby accurately and efficiently intervening in faults in advance and ensuring the normal execution of distributed cluster tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cluster fault handling technology, and in particular to a proactive fault defense method, apparatus, device, and storage medium. Background Technology

[0002] Currently, due to the surge in task processing volume, the introduction of clustered distributed systems has significantly improved task processing efficiency. However, if any node in a clustered distributed system fails during task execution, it may cause the task execution to be interrupted. Therefore, timely fault handling is of great value.

[0003] However, in current technologies, when handling cluster faults, fault signals can usually only be detected at the hardware level, while software error signals that may cause hardware faults cannot be captured. Therefore, the perception dimension is missing, and the best time window for early intervention is missed. When handling faults, the analysis and processing are carried out independently from a single computing node, ignoring the fault propagation effect between nodes, which reduces the accuracy of fault identification. After the fault is identified, the defense strategy is limited to resource reservation and task rescheduling. Summary of the Invention

[0004] This invention provides a proactive fault defense method, apparatus, device, and storage medium to achieve precise fault defense.

[0005] According to a first aspect of the present invention, a proactive fault defense method is provided, the method comprising: constructing a dynamic fault propagation map based on the computing node cluster under normal operating conditions, wherein a node agent is deployed on each computing node of the node cluster. When a soft error signal is received from a node agent, the dynamic fault propagation map is updated by combining the link status signal reported by the network sensing agent. Predict the risk level of each computing node based on the updated dynamic fault propagation graph; For each computing node, a defense strategy matching the risk level is determined, and the defense strategy is executed to perform fault defense on the computing node.

[0006] According to another aspect of the present invention, an active fault defense device is provided, the device comprising: a dynamic fault propagation graph construction module, configured to construct a dynamic fault propagation graph based on the computing node cluster under normal operating conditions, wherein a node agent is deployed on each computing node of the node cluster. The dynamic fault propagation map update module is used to update the dynamic fault propagation map by combining the link status signal reported by the network sensing agent when a soft error signal reported by the node agent is received. The risk level prediction module is used to predict the risk level of each computing node based on the updated dynamic fault propagation map. The fault defense module is used to determine a defense strategy that matches the risk level for each computing node, and execute the defense strategy to perform fault defense on the computing node.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: one or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any embodiment of the present invention.

[0008] According to another aspect of the present invention, a storage medium for computer-executable instructions is provided, on which a computer program is stored, which, when executed by a processor, implements the method as described in any of the embodiments of the present invention.

[0009] The technical solution of this invention constructs and updates a dynamic fault propagation map, deeply integrates hardware soft error signals with communication topology, predicts the risk level of each computing node in the node cluster based on the propagation path of the fault in the dynamic fault propagation map, and adopts a defense strategy that matches the risk level to perform differentiated fault defense, thereby accurately and efficiently intervening in faults in advance and ensuring the normal execution of distributed cluster tasks.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart of an active fault defense method provided according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram illustrating the execution process of the logical topology bypass strategy provided in Embodiment 1 of the present invention; Figure 3 This is a flowchart of another active fault defense method provided according to Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the structure of an active fault defense device provided according to Embodiment 3 of the present invention; Figure 5 This is a structural block diagram of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation

[0013] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0014] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or terminal device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or terminal devices.

[0015] Example 1 Figure 1 This is a flowchart of a proactive fault defense method provided in Embodiment 1 of the present invention. This embodiment is applicable to the proactive defense of faults in distributed clusters. The method can be executed by a proactive fault defense device, which can be implemented in hardware and / or software, and can be integrated into an electronic device with data processing capabilities. Figure 1 As shown, the method includes: S101, under normal operating conditions of the computing node cluster, constructs a dynamic fault propagation map based on the computing node cluster.

[0016] Optionally, a dynamic fault propagation graph is constructed based on the computing node cluster, including: determining each graph node based on the computing nodes contained in the computing node cluster, and using the local health of the computing nodes as the node weight of the corresponding graph node; determining graph edges based on the initial communication dependencies of each computing node in the computing node cluster, and setting the edge weight of each graph edge to a preset value, wherein the edge weight is used to represent the probability of a fault propagating along the communication link corresponding to the graph edge; constructing a directed weighted graph based on the graph nodes, graph edges, and edge weights, and using the directed weighted graph as the dynamic fault propagation graph.

[0017] Specifically, the computing node cluster contains multiple computing nodes with communication relationships, and each computing node is equipped with a node agent. Each node agent collects multi-dimensional soft error signals from its assigned computing node. Additionally, outside the computing node cluster, there is a node deploying a network-aware agent. This network-aware agent primarily monitors the link quality of each computing node in the cluster to identify abnormal propagation at the communication layer. Both the node agents and the network-aware agent report their collected data to a global agent located in the management center.

[0018] During normal operation of the computing node cluster, the global agent determines each graph node V based on the computing nodes contained in the cluster, and uses the local health of each computing node as the node weight of the corresponding graph node. Furthermore, since the connection relationships of each computing node are noted in the global agent during the cluster creation phase, the global agent can determine graph edges E based on the initial communication dependencies of each computing node, such as neighbor relationships or connection relationships. Additionally, in the initial stage, the edge weight W of each graph edge is set to a preset value. The edge weight refers to the probability that a fault will propagate along the communication link corresponding to that graph edge. This preset value can be derived from historical data statistics. Based on the above information, the global agent constructs a directed weighted graph G=(V,E,W) and uses the obtained directed weighted graph G as the dynamic fault propagation graph.

[0019] S102, when a soft error signal is received from a node agent, the dynamic fault propagation map is updated by combining the link status signal reported by the network sensing agent.

[0020] Optionally, when a soft error signal is received from a node agent, the dynamic fault propagation graph is updated in conjunction with the link status signal reported by the network perception agent. This includes: identifying abnormal computing nodes based on the reported soft error signal, and identifying target graph nodes matching the abnormal computing nodes in the dynamic fault propagation graph. The soft error signal includes the growth rate of single-bit error correction code errors in the memory layer, the long-tail distribution of the execution time of specified operators in the computing layer, and the number of remote direct memory access retransmission packets and the frequency of negative acknowledgment signals belonging to the network layer of the computing node; determining the latest health status of the abnormal computing node based on the soft error signal, and updating the node weights of the target graph nodes based on the latest health status; and updating the graph edges associated with the target graph nodes and their edge weights based on the reported link status signal. The link status signal includes link latency jitter, bandwidth throughput change, forward error correction code rate, and explicit congestion notification surge signal.

[0021] Specifically, the node agents deployed on each computing node will monitor the corresponding computing nodes in real time through underlying system tools. Specifically, they can monitor the hardware counters within the computing node used to record the number of times specific events occur, and report the collected soft error signals to the global agent. The soft error signals include the growth rate of the number of single-bit error correction codes in the memory layer, the long-tail distribution of the execution time of specified operators in the computing layer, and the number of Remote Direct Memory Access (RDMA) retransmission packets and the frequency of negative acknowledgment (NACK) signals belonging to the network layer of the computing node. Among them, the single-bit error correction code error correction rate refers to the speed or magnitude of the increase in the number of times the hardware automatically corrects single-bit errors per unit time. When this indicator rises abnormally, it indicates that the video memory or main memory of the computing node is experiencing a failure rate exceeding the normal range, usually indicating that a more serious hardware failure is about to occur. The long-tail distribution of execution time refers to the fact that in a large number of identical computing tasks, the execution time of most tasks is very short and stable, but the execution time of a very small number of tasks is abnormally long, causing a long "tail" to appear on the right side of the overall time distribution graph. When the node agent detects a significant long-tail distribution, it indicates that the GPU in the computing node may be experiencing local overheating leading to frequency reduction, video memory bandwidth being preempted by other processes, internal bus conflicts, or unstable power supply. The NACK signal frequency refers to the number of times a NACK signal is received per unit time. The higher this frequency, the more severe the packet loss or data corruption in the network link. The RDMA retransmission packet count refers to the number of packets the sender receives when it receives a NACK signal. The number of data packets retransmitted after a signal or after the ACK (acknowledgment) timeout is a direct consequence of the NACK signal. A surge in the number of retransmitted packets directly reflects problems such as network link congestion, optical module failure, or physical link quality degradation.

[0022] The global intelligent machine (GAM) identifies abnormal computing nodes based on received soft error signals and determines target graph nodes matching these abnormal nodes in a dynamic fault map. It then determines the latest health status of the abnormal computing nodes based on the acquired soft error signals. This health status can be obtained by scoring various soft error signals using preset scoring rules or by calculation using a specified method. Therefore, soft error signals represent the underlying symptoms of the computing nodes, while the latest health status is a diagnostic score obtained by the global intelligent machine after comprehensively evaluating these symptoms. In this embodiment, the node weights of the target graph nodes are updated based on the calculated latest health status. During the update, the latest health status can be used directly as the node weight, or the result of processing the latest health status according to a preset algorithm can be used as the node weight. Therefore, the node weight can be the latest health status or an intermediate parameter obtained through calculation based on the latest health status.

[0023] Furthermore, this embodiment also updates the graph edges associated with the target graph node based on the link status signals reported by the network sensing agent. For example, when a new graph node has a communication association with the target graph node, the number of graph edges associated with the target graph node will increase; when the communication association between the original graph node and the target graph node is interrupted, the number of graph edges associated with the target graph node will decrease. Additionally, the link status signals include information such as link latency jitter, bandwidth throughput variation, forward error correction code rate, and explicit congestion notification surge signals. Link latency jitter refers to the inconsistent network latency experienced by different data packets in the same service flow, reflecting the stability of network transmission; bandwidth throughput variation refers to the fluctuation range of actual throughput; forward error correction code rate reflects the anti-interference capability and reliability of the link; and explicit congestion notification surge signals indicate that intermediate nodes (such as switches and routers) in the network path are facing a serious buffer overflow risk. Based on the link state signals mentioned above, the latest fault propagation probability of each graph edge can be determined, and the weights of the graph edges can be updated accordingly based on the latest determined fault propagation probability.

[0024] S103, predict the risk level of each computing node based on the updated dynamic fault propagation map.

[0025] Optionally, the risk level of each computing node is predicted based on the updated dynamic fault propagation graph. This includes extracting node weights, a first comprehensive risk index for each adjacent graph node, and edge weights of the graph edges connected to adjacent graph nodes for each graph node in the updated dynamic fault propagation graph. A second comprehensive risk index is determined for the computing node corresponding to the graph node by multiplying the node weight by a preset first weight coefficient, and then adding the sum of the products of the first comprehensive risk index of each adjacent graph node, the corresponding edge weight, and a preset second weight coefficient. The risk level of the computing node is determined based on the second comprehensive risk index, where the risk level includes normal operation, presence of soft errors, and impending hard failure.

[0026] Optionally, the method also includes: for each computing node, extracting edge weights and historical fault data associated with the computing node from the updated dynamic fault propagation graph; and estimating the remaining safe runtime of the computing node based on the risk level, edge weights, and historical fault data.

[0027] Specifically, the comprehensive risk coefficient of each computing node can be predicted based on the updated dynamic fault propagation map. For example, the corresponding comprehensive risk index for each map node is calculated using the following formula (1): (1) in, Indicates the first Graph nodes Indicates the relationship with the first Each graph node is connected to its adjacent graph nodes via graph edges, and the number of adjacent graph nodes is at least one. Indicates the first The second comprehensive risk index for each graph node. Indicates the first i The node weights of each graph node Indicates the first The first comprehensive risk index of adjacent graph nodes. Indicates the first The graph node and the first The edge weights of graph edges between adjacent graph nodes. This represents the first weighting coefficient. This represents the second weighting coefficient. In this implementation, the second comprehensive risk index determined above will be used to determine the... The risk level of each graph node can be determined by, for example, a first threshold and a second threshold. When the second comprehensive risk index is less than the first threshold, the risk level is determined to be normal operation. When it is between the first threshold and the second threshold, the risk level is determined to be soft error. When it is greater than the second threshold, the risk level is determined to be hard failure. The higher the risk level, the better.

[0028] Furthermore, this implementation determines the remaining safe runtime of each computing node based on the calculated risk level and dynamic fault propagation graph. Specifically, it extracts the edge weights associated with the computing nodes and historical fault data from the dynamic fault propagation graph. The edge weights represent the probability of a fault propagating along the communication link. If a computing node has a large edge weight, it indicates that it is not only prone to problems itself, but also very likely to propagate faults to adjacent computing nodes, dragging down the entire cluster. Such computing nodes need to be dealt with more urgently, and their remaining safe runtime will be judged to be shorter in the system evaluation. Historical fault data contains the frequency and pattern of past faults of computing nodes. If a computing node has a large number of historical faults, it indicates that its hardware is severely aging or has design flaws, and the system's trust in its future will decrease, thereby shortening the estimated remaining safe runtime. The aforementioned risk level reflects the health status of the computing node. The higher the risk level, the more serious the current hardware or link problems of the computing node, and the greater the risk of complete downtime. In this embodiment, the remaining safe runtime of the computing node is calculated based on information from three dimensions: risk level, edge weight, and historical failure data. For example, it can safely run for another hour, thereby predicting the operating status of the computing node in advance.

[0029] S104. Determine a defense strategy that matches the risk level for each computing node, and execute the defense strategy to defend against faults in the computing nodes.

[0030] Optionally, a defense strategy matching the risk level is determined for each computing node, and the defense strategy is executed to perform fault defense on the computing node, including: when the risk level of the computing node is that there is a soft error, the task type of the task running on the computing node is identified, wherein the task type includes communication-intensive tasks or state-sensitive tasks; when the task type is a communication-intensive task, the defense strategy is determined to be a logical topology bypass strategy, and the computing node is switched from the critical communication path to the bypass path without interrupting the task according to the logical topology bypass strategy; when the task type is a state-sensitive task, the defense strategy is determined to be a hot migration strategy, and the start time of the incremental checkpoint is determined according to the hot migration strategy based on the remaining safe runtime. When the start time is reached, a backup healthy computing node is started, the node agent on the abnormal computing node is instructed to start the incremental checkpoint, the parameter changes since the last global snapshot time are saved as incremental data, and the data is migrated to the backup healthy computing node through background asynchronous transmission. When the task execution is determined to be finished, the task execution context is switched to the backup healthy computing node.

[0031] Optionally, a defense strategy matching the risk level is determined for each computing node, and the defense strategy is executed to defend against faults in the computing node. This includes: when the risk level of a computing node is that a hard fault is about to occur, identifying the status of idle nodes in the computing node cluster, where the status includes whether there are idle nodes or not; when there are idle nodes, the defense strategy is determined to be a hot migration strategy, and the start time of the incremental checkpoint is determined according to the remaining safe runtime according to the hot migration strategy. When the start time is reached, a backup healthy computing node is started, and the node agent on the abnormal computing node is instructed to start the incremental checkpoint, saving the parameter changes since the last global snapshot as incremental data, and migrating it to the backup healthy computing node through background asynchronous transmission. When the task execution is determined to be finished, the task execution context is switched to the backup healthy computing node; when there are no idle nodes, the defense strategy is determined to be an active degradation strategy, and the computing frequency of the computing node is reduced or the upper limit of video memory usage is limited according to the active degradation strategy.

[0032] Specifically, in this embodiment, after predicting the risk level, a defense strategy corresponding to the risk level can be adopted to defend against faults in the computing node. For the risk level of soft faults, the corresponding defense strategy will be further determined based on the task type on the computing node. For example, when it is a communication-intensive task, the defense strategy is determined to be a logical topology bypass strategy, and when it is a state-sensitive task, the defense strategy is determined to be a hot migration strategy. For the risk level of an impending hard fault, the corresponding defense strategy will be further determined based on the status of idle nodes. For example, when there are idle nodes, the defense strategy is determined to be a hot migration strategy, and when there are no idle nodes, the defense strategy is determined to be an active degradation strategy. The specific execution process of the above three defense strategies will be explained in detail below.

[0033] Among them, such as Figure 2 The diagram illustrates the execution process of the logical topology bypass strategy. The network-aware agent notifies the global agent to dynamically modify the logical connections of the communication ring or tree without interrupting tasks. For example, if computing node 3 is determined to be an abnormal computing node, but the communication ring containing node 3 requires a large amount of communication transmission, and the presence of node 3 affects normal communication, then node 3 will be switched from the critical communication path to bypass, allowing it to only undertake computing tasks or communicate through a low-speed backup link, and no longer undertake critical communication tasks. Furthermore, the proactive degradation strategy refers to the situation where a computing node is determined to be about to experience a hard failure, but there are no replaceable healthy computing nodes in the computing node cluster. To ensure normal task execution, the node agent will automatically reduce the computing frequency of the computing node or limit the upper limit of its memory usage to extend its lifespan. This sacrifices some performance for stability, preventing immediate crashes and buying time for migration.

[0034] The hot migration strategy refers to a mechanism where, when a compute node is about to experience a hard failure and there is a replaceable healthy compute node in the cluster, or when a compute node has a soft error and is executing a state-sensitive task, the global agent marks the compute node running the new task as unschedulable. Based on the remaining safe runtime, the start time of the incremental checkpoint is determined according to the hot migration strategy. For example, if the current time is 10:00 and the remaining safe runtime is 1 hour, the start time of the incremental checkpoint can be determined to be 10:30. Of course, other times are also possible, but the start time cannot exceed 11:00. Upon reaching the start time, a backup compute node is activated, and the node agent on the faulty compute node initiates the incremental checkpoint. The incremental checkpoint only saves the parameter changes since the last global snapshot as incremental data. Because model parameters are very large during AI training, saving the entire dataset each time would be too time-consuming. Incremental checkpointing only records and saves the data that has changed since the last backup, which greatly reduces the amount of data that needs to be processed and shortens the state saving time. While the abnormal computing node continues to execute its computational tasks, the system utilizes idle network bandwidth to asynchronously send this incremental data to the backup healthy computing node in the background. Because it is asynchronous, it does not block the normal computation of the current task. In addition, the global intelligence precisely calculates and waits for the current iteration task to complete, and then immediately and seamlessly switches control of the task, i.e., the execution context, to the backup healthy computing node that has already received the incremental data. Because the switch occurs in the gap between task completions, the front-end business and user experience do not experience any interruption or delay. Thus, through the method of "incremental saving + background transfer + instant switch", critical tasks can be safely transferred with minimal performance loss before the hardware completely fails, thereby ensuring the continuity and high availability of the entire distributed cluster.

[0035] The technical solution of this invention constructs and updates a dynamic fault propagation map, deeply integrates hardware soft error signals with communication topology, predicts the risk level of each computing node in the node cluster based on the propagation path of the fault in the dynamic fault propagation map, and adopts a defense strategy that matches the risk level to perform differentiated fault defense, thereby accurately and efficiently intervening in faults in advance and ensuring the normal execution of distributed cluster tasks.

[0036] Example 2 Figure 3This is a flowchart of another proactive fault defense method provided by an embodiment of the present invention. Based on the above embodiments, after executing the defense strategy to defend against faults in the computing nodes, this embodiment further includes: obtaining actual fault information of the computing nodes and obtaining the matching degree between the actual fault information and the risk level; adjusting the dynamic fault propagation map, the extraction threshold corresponding to the soft error signal, the initiation advance of the incremental checkpoint, the first weight coefficient, and the second weight coefficient according to the matching degree, wherein each node agent reports the software error signal based on the extraction threshold, such as... Figure 3 As shown, the method includes: S201, under normal operating conditions of the computing node cluster, constructs a dynamic fault propagation map based on the computing node cluster.

[0037] Optionally, a dynamic fault propagation graph is constructed based on the computing node cluster, including: determining each graph node based on the computing nodes contained in the computing node cluster, and using the local health of the computing nodes as the node weight of the corresponding graph node; determining graph edges based on the initial communication dependencies of each computing node in the computing node cluster, and setting the edge weight of each graph edge to a preset value, wherein the edge weight is used to represent the probability of a fault propagating along the communication link corresponding to the graph edge; constructing a directed weighted graph based on the graph nodes, graph edges, and edge weights, and using the directed weighted graph as the dynamic fault propagation graph.

[0038] S202, when a soft error signal is received from a node agent, the dynamic fault propagation map is updated by combining the link status signal reported by the network sensing agent.

[0039] Optionally, when a soft error signal is received from a node agent, the dynamic fault propagation graph is updated in conjunction with the link status signal reported by the network perception agent. This includes: identifying abnormal computing nodes based on the reported soft error signal, and identifying target graph nodes matching the abnormal computing nodes in the dynamic fault propagation graph. The soft error signal includes the growth rate of single-bit error correction code errors in the memory layer, the long-tail distribution of the execution time of specified operators in the computing layer, and the number of remote direct memory access retransmission packets and the frequency of negative acknowledgment signals belonging to the network layer of the computing node; determining the latest health status of the abnormal computing node based on the soft error signal, and updating the node weights of the target graph nodes based on the latest health status; and updating the graph edges associated with the target graph nodes and their edge weights based on the reported link status signal. The link status signal includes link latency jitter, bandwidth throughput change, forward error correction code rate, and explicit congestion notification surge signal.

[0040] S203, predict the risk level of each computing node based on the updated dynamic fault propagation map.

[0041] Optionally, the risk level of each computing node is predicted based on the updated dynamic fault propagation graph. This includes extracting node weights, a first comprehensive risk index for each adjacent graph node, and edge weights of the graph edges connected to adjacent graph nodes for each graph node in the updated dynamic fault propagation graph. A second comprehensive risk index is determined for the computing node corresponding to the graph node by multiplying the node weight by a preset first weight coefficient, and then adding the sum of the products of the first comprehensive risk index of each adjacent graph node, the corresponding edge weight, and a preset second weight coefficient. The risk level of the computing node is determined based on the second comprehensive risk index, where the risk level includes normal operation, presence of soft errors, and impending hard failure.

[0042] Optionally, the method also includes: for each computing node, extracting edge weights and historical fault data associated with the computing node from the updated dynamic fault propagation graph; and estimating the remaining safe runtime of the computing node based on the risk level, edge weights, and historical fault data.

[0043] S204. Determine a defense strategy that matches the risk level for each compute node, and execute the defense strategy to defend against faults in the compute nodes.

[0044] Optionally, a defense strategy matching the risk level is determined for each computing node, and the defense strategy is executed to perform fault defense on the computing node, including: when the risk level of the computing node is that there is a soft error, the task type of the task running on the computing node is identified, wherein the task type includes communication-intensive tasks or state-sensitive tasks; when the task type is a communication-intensive task, the defense strategy is determined to be a logical topology bypass strategy, and the computing node is switched from the critical communication path to the bypass path without interrupting the task according to the logical topology bypass strategy; when the task type is a state-sensitive task, the defense strategy is determined to be a hot migration strategy, and the start time of the incremental checkpoint is determined according to the hot migration strategy based on the remaining safe runtime. When the start time is reached, a backup healthy computing node is started, the node agent on the abnormal computing node is instructed to start the incremental checkpoint, the parameter changes since the last global snapshot time are saved as incremental data, and the data is migrated to the backup healthy computing node through background asynchronous transmission. When the task execution is determined to be finished, the task execution context is switched to the backup healthy computing node.

[0045] Optionally, a defense strategy matching the risk level is determined for each computing node, and the defense strategy is executed to defend against faults in the computing node. This includes: when the risk level of a computing node is that a hard fault is about to occur, identifying the status of idle nodes in the computing node cluster, where the status includes whether there are idle nodes or not; when there are idle nodes, the defense strategy is determined to be a hot migration strategy, and the start time of the incremental checkpoint is determined according to the remaining safe runtime according to the hot migration strategy. When the start time is reached, a backup healthy computing node is started, and the node agent on the abnormal computing node is instructed to start the incremental checkpoint, saving the parameter changes since the last global snapshot as incremental data, and migrating it to the backup healthy computing node through background asynchronous transmission. When the task execution is determined to be finished, the task execution context is switched to the backup healthy computing node; when there are no idle nodes, the defense strategy is determined to be an active degradation strategy, and the computing frequency of the computing node is reduced or the upper limit of video memory usage is limited according to the active degradation strategy.

[0046] S205: Obtain the actual fault information of the computing node and the matching degree between the actual fault information and the risk level. Adjust the dynamic fault propagation map, the extraction threshold corresponding to the soft error signal, the start advance of the incremental checkpoint, the first weight coefficient and the second weight coefficient according to the matching degree.

[0047] Specifically, in this embodiment, when the global agent determines that a certain computing node is at risk, such as determining that a hardware failure is about to occur, and takes corresponding defense strategies, the global agent will continuously track the actual failure information of the computing node, compare the predicted risk level with the actual failure to obtain the matching degree, and adjust the relevant parameters involved in the risk level prediction process based on the obtained matching degree, such as the dynamic fault propagation map, the extraction threshold corresponding to the soft error signal, the start advance of the incremental checkpoint, the first weight coefficient, and the second weight coefficient.

[0048] In this scenario, a low matching degree indicates inaccurate prediction results. In such cases, the structure or edge weights of the dynamic fault propagation graph are modified to better reflect the actual physical network. If normal fluctuations are misjudged as soft errors, the alarm threshold is raised because each node agent reports software error signals based on extracted thresholds. Conversely, if potential problems are frequently missed, the threshold is lowered to make the system more sensitive. If each hot migration results in task interruption due to inaccurate timing, the checkpoint triggering time is adjusted to ensure sufficient time for data transfer next time. The first and second weighting coefficients are used to calculate the comprehensive risk index. The global agent adjusts their weights based on the review results; for example, if the influence of adjacent computing nodes is found to be more significant than its own state, the corresponding weighting coefficient is increased. This implementation, through a closed loop of prediction, execution, review, and adjustment, enables the global agent to possess machine learning and adaptive capabilities. It accumulates experience as the cluster runs, leading to more accurate fault predictions, better defense strategies, and smoother hot migrations in the future.

[0049] It is worth mentioning that this implementation continuously collects soft error signals through node agents and uses a dynamic fault propagation graph to predict the probability of fault occurrence and propagation path. This proactive intervention during the early warning period before a fault actually occurs represents a paradigm shift from passive response to proactive prediction. By constructing and updating the dynamic fault propagation graph, and modeling the health of computing nodes in conjunction with the communication dependencies between tasks, the propagation risk of an anomaly in a particular node to its neighboring nodes can be quantified, achieving a leap from single-point monitoring to global correlation analysis. Through dynamic logical topology reconstruction technology, the logical path of the communication loop is modified in real time without interrupting training, temporarily bypassing high-risk nodes, thus solving the problem of single-point jitter dragging down the entire system.

[0050] The technical solution of this invention constructs and updates a dynamic fault propagation map, deeply integrates hardware soft error signals with communication topology, predicts the risk level of each computing node in the node cluster based on the propagation path of the fault in the dynamic fault propagation map, and adopts a defense strategy that matches the risk level to perform differentiated fault defense, thereby accurately and efficiently intervening in faults in advance and ensuring the normal execution of distributed cluster tasks.

[0051] Example 3 Figure 4 This is a schematic diagram of the structure of an active fault defense device provided in an embodiment of the present invention. Figure 4 As shown, the device includes: a dynamic fault propagation map construction module 310, a dynamic fault propagation map update module 320, a risk level prediction module 330, and a fault defense module 340.

[0052] Among them, the dynamic fault propagation graph construction module 310 is used to construct a dynamic fault propagation graph based on the computing node cluster under normal operating conditions. In this module, a node agent is deployed on each computing node of the node cluster. The dynamic fault propagation map update module 320 is used to update the dynamic fault propagation map by combining the link status signal reported by the network sensing agent when a soft error signal reported by the node agent is received. Risk level prediction module 330 is used to predict the risk level of each computing node based on the updated dynamic fault propagation map. The fault defense module 340 is used to determine a defense strategy that matches the risk level for each computing node and execute the defense strategy to perform fault defense on the computing node.

[0053] Optionally, the dynamic fault propagation graph construction module 310 is used to determine each graph node based on the computing nodes contained in the computing node cluster, and to use the local health of the computing nodes as the node weight of the corresponding graph node. The graph edges are determined based on the initial communication dependencies of each computing node in the computing node cluster, and the edge weights of each graph edge are set to preset values. The edge weights are used to represent the probability that a fault will propagate along the communication link corresponding to the graph edge. A directed weighted graph is constructed based on the graph nodes, graph edges, and edge weights, and the directed weighted graph is used as a dynamic fault propagation graph.

[0054] Optionally, the dynamic fault propagation map update module 320 is used to determine the abnormal computing node based on the reported soft error signal, and to determine the target map node that matches the abnormal computing node in the dynamic fault propagation map. The soft error signal includes the growth rate of the number of single-bit error correction codes in the memory layer, the long-tail distribution of the execution time of the specified operator in the computing layer, and the number of remote direct memory access retransmission packets and the frequency of negative acknowledgment signals belonging to the network layer of the computing node. The latest health status of the abnormal computing nodes is determined based on the soft error signal, and the node weights of the target graph nodes are updated based on the latest health status. The graph edges associated with the target graph node and their weights are updated based on the reported link status signals. The link status signals include link latency jitter, bandwidth throughput change, forward error correction code rate, and explicit congestion notification surge signal.

[0055] Optionally, the risk level prediction module 330 is used to extract the node weight, the first comprehensive risk index of each adjacent node, and the edge weight of the graph edge connected to the adjacent node for each graph node in the updated dynamic fault propagation graph. The second comprehensive risk index of the computation node corresponding to the graph node is determined by multiplying the node weight by the preset first weight coefficient, and adding the sum of the product of the first comprehensive risk index of each adjacent graph node and the corresponding edge weight and the preset second weight coefficient. The risk level of the computing node is determined based on the second comprehensive risk index, which includes normal operation, presence of soft errors, and impending hard failure.

[0056] Optionally, the device also includes a remaining safe runtime estimation module, which is used to extract the edge weights and historical fault data associated with the computing nodes from the updated dynamic fault propagation graph for each computing node. The remaining safe runtime of a node is estimated based on risk level, edge weight, and historical failure data.

[0057] Optionally, the fault prevention module 340 is used to identify the task type of the task running on the computing node when the risk level of the computing node is that a soft error exists. The task type includes communication-intensive tasks or state-sensitive tasks. When the task type is a communication-intensive task, the defense strategy is determined to be a logical topology bypass strategy, and the computing node is switched from the critical communication path to the bypass path without interrupting the task, in accordance with the logical topology bypass strategy. When the task type is a state-sensitive task, the defense strategy is determined to be a hot migration strategy. According to the hot migration strategy, the start time of the incremental checkpoint is determined based on the remaining safe runtime. When the start time is reached, the backup healthy computing node is started, and the node agent on the abnormal computing node is instructed to start the incremental checkpoint. The parameter changes since the last global snapshot time are saved as incremental data and migrated to the backup healthy computing node through background asynchronous transmission. When the task execution is determined to be finished, the task execution context is switched to the backup healthy computing node.

[0058] Optionally, the fault defense module 340 is used to identify the status of idle nodes in the computing node cluster when the risk level of the computing node is about to occur a hard failure, wherein the status includes whether there are idle nodes or not. When the state indicates that there are idle nodes, the defense strategy is determined to be a hot migration strategy. According to the hot migration strategy, the start time of the incremental checkpoint is determined based on the remaining safe runtime. When the start time is reached, the backup healthy computing node is started, and the node agent on the abnormal computing node is instructed to start the incremental checkpoint. The parameter changes since the last global snapshot are saved as incremental data and migrated to the backup healthy computing node through background asynchronous transmission. When the task execution is determined to be finished, the task execution context is switched to the backup healthy computing node. When there are no idle nodes, the defense strategy is determined to be an active degradation strategy, and the computing frequency of the computing nodes is reduced or the upper limit of video memory usage is limited in accordance with the active degradation strategy.

[0059] Optionally, the device also includes an adjustment module for obtaining actual fault information of computing nodes and obtaining the matching degree between actual fault information and risk level; The dynamic fault propagation map, the extraction threshold corresponding to the soft error signal, the start advance of the incremental checkpoint, the first weight coefficient and the second weight coefficient are adjusted according to the matching degree. Among them, each node agent reports the software error signal based on the extraction threshold.

[0060] The active fault defense device provided in this embodiment of the invention can execute an active fault defense method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0061] Example 4 Figure 5 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0062] The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.

[0063] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0064] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other electronic devices through computer networks such as the Internet and / or various telecommunications networks.

[0065] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as proactive fault defense methods.

[0066] That is, under normal operating conditions of the computing node cluster, a dynamic fault propagation map is constructed based on the computing node cluster, wherein a node agent is deployed on each computing node of the node cluster. When a soft error signal is received from a node agent, the dynamic fault propagation map is updated by combining the link status signal reported by the network sensing agent. Predict the risk level of each computing node based on the updated dynamic fault propagation graph; For each computing node, a defense strategy matching the risk level is determined, and the defense strategy is executed to defend against faults in the computing nodes.

[0067] In some embodiments, the proactive fault defense method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the proactive fault defense method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the proactive fault defense method by any other suitable means (e.g., by means of firmware).

[0068] Various embodiments of the apparatuses and techniques described above herein can be implemented in digital electronic circuit devices, integrated circuit devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), device-on-a-chip (SoC) devices, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable device including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage device, at least one input device, and at least one output device, and transmitting data and instructions to the storage device, the at least one input device, and the at least one output device.

[0069] Computer programs used to implement the proactive fault defense method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer or a special-purpose computer, such that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, or as a standalone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0070] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution apparatus, device, or electronic device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage electronics, magnetic storage electronics, or any suitable combination thereof.

[0071] To provide interaction with a user, the devices and techniques described herein can be implemented on an electronic device having: a display device (e.g., a touchscreen) for displaying information to the user; and buttons through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0072] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0073] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A proactive fault defense method, characterized in that, Applied to a global intelligent agent, the method includes: Under normal operating conditions of the computing node cluster, a dynamic fault propagation graph is constructed based on the computing node cluster, wherein a node agent is deployed on each computing node of the node cluster. When a soft error signal is received from a node agent, the dynamic fault propagation map is updated by combining the link status signal reported by the network sensing agent. Predict the risk level of each computing node based on the updated dynamic fault propagation graph; For each computing node, a defense strategy matching the risk level is determined, and the defense strategy is executed to perform fault defense on the computing node.

2. The method according to claim 1, characterized in that, The step of constructing a dynamic fault propagation graph based on the computing node cluster includes: Each graph node is determined based on the computing nodes contained in the computing node cluster, and the local health of the computing node is used as the node weight of the corresponding graph node. The graph edges are determined based on the initial communication dependencies of each computing node in the computing node cluster, and the edge weights of each graph edge are set to preset values. The edge weights are used to represent the probability that a fault will propagate along the communication link corresponding to the graph edge. A directed weighted graph is constructed based on the graph nodes, graph edges, and edge weights, and the directed weighted graph is used as the dynamic fault propagation graph.

3. The method according to claim 1, characterized in that, When a soft error signal is received from a node agent, the dynamic fault propagation map is updated by combining the link status signal reported by the network sensing agent, including: The abnormal computing node is determined based on the reported soft error signal, and the target graph node matching the abnormal computing node is determined in the dynamic fault propagation graph. The soft error signal includes the growth rate of the number of single-bit error correction codes in the memory layer, the long-tail distribution of the execution time of the specified operator in the computing layer, and the number of remote direct memory access retransmission packets and the frequency of negative acknowledgment signals belonging to the network layer of the computing node. The latest health status of the abnormal computing node is determined based on the soft error signal, and the node weight of the target graph node is updated based on the latest health status. The graph edges associated with the target graph node and their weights are updated based on the reported link status signals. The link status signals include link latency jitter, bandwidth throughput change, forward error correction code rate, and explicit congestion notification surge signal.

4. The method according to claim 3, characterized in that, The step of predicting the risk level of each computing node based on the updated dynamic fault propagation map includes: For each graph node in the updated dynamic fault propagation graph, extract the node weight, the first comprehensive risk index of each adjacent graph node, and the edge weight of the graph edge connected to the adjacent graph node. The second comprehensive risk index of the computing node corresponding to the graph node is determined by multiplying the node weight by the preset first weight coefficient, and adding the sum of the products of the first comprehensive risk index of each adjacent graph node, the corresponding edge weight, and the preset second weight coefficient. The risk level of the computing node is determined based on the second comprehensive risk index, wherein the risk level includes normal operation, presence of soft errors, and impending hard failure.

5. The method according to claim 1, characterized in that, The method further includes: For each computing node, the edge weights and historical fault data associated with the computing node are extracted from the updated dynamic fault propagation graph; The remaining safe runtime of the computing node is estimated based on the risk level, the edge weight, and the historical failure data.

6. The method according to claim 5, characterized in that, The step of determining a defense strategy matching the risk level for each computing node and executing the defense strategy to perform fault defense on the computing node includes: When the risk level of a computing node is that a soft error exists, the task type of the task running on the computing node is identified, wherein the task type includes communication-intensive tasks or state-sensitive tasks. When the task type is the communication-intensive task, the defense strategy is determined to be the logical topology bypass strategy, and the computing node is switched from the critical communication path to the bypass path without interrupting the task, in accordance with the logical topology bypass strategy. When the task type is a state-sensitive task, the defense strategy is determined to be a hot migration strategy. According to the hot migration strategy, the start time of the incremental checkpoint is determined based on the remaining safe runtime. When the start time is reached, the backup healthy computing node is started, and the node agent on the abnormal computing node is instructed to start the incremental checkpoint. The parameter changes since the last global snapshot are saved as incremental data and migrated to the backup healthy computing node through background asynchronous transmission. When the task execution is determined to be finished, the task execution context is switched to the backup healthy computing node.

7. The method according to claim 5, characterized in that, The step of determining a defense strategy matching the risk level for each computing node and executing the defense strategy to perform fault defense on the computing node includes: When the risk level of a compute node is that a hard failure is about to occur, the status of idle nodes in the compute node cluster is identified, wherein the status includes the presence of idle nodes or the absence of idle nodes; When the state is that there are idle nodes, the defense strategy is determined to be a hot migration strategy. According to the hot migration strategy, the start time of the incremental checkpoint is determined based on the remaining safe runtime. When the start time is reached, the backup healthy computing node is started, and the node agent on the abnormal computing node is instructed to start the incremental checkpoint. The parameter changes since the last global snapshot time are saved as incremental data and migrated to the backup healthy computing node through background asynchronous transmission. When the task execution is determined to be finished, the task execution context is switched to the backup healthy computing node. When the state is that there are no idle nodes, the defense strategy is determined to be an active degradation strategy, and the computing frequency of the computing node is reduced or the upper limit of video memory usage is limited in accordance with the active degradation strategy.

8. The method according to claim 4, characterized in that, After executing the defense strategy to perform fault defense on the computing node, the method further includes: Obtain actual fault information of computing nodes, and obtain the matching degree between the actual fault information and the risk level; The dynamic fault propagation map, the extraction threshold corresponding to the soft error signal, the start advance of the incremental checkpoint, the first weight coefficient, and the second weight coefficient are adjusted according to the matching degree. Each node agent reports the software error signal based on the extraction threshold.

9. A fault-prevention device, characterized in that, The device includes: The dynamic fault propagation graph construction module is used to construct a dynamic fault propagation graph based on the computing node cluster under normal operating conditions, wherein a node agent is deployed on each computing node of the node cluster. The dynamic fault propagation map update module is used to update the dynamic fault propagation map by combining the link status signal reported by the network sensing agent when a soft error signal reported by the node agent is received. The risk level prediction module is used to predict the risk level of each computing node based on the updated dynamic fault propagation map. The fault defense module is used to determine a defense strategy that matches the risk level for each computing node, and execute the defense strategy to perform fault defense on the computing node.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

11. A storage medium for computer-executable instructions, wherein a computer program is stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.