Model training node fault recovery method, fault detection equipment, device and medium

By automatically detecting and recovering node failures in distributed computing clusters through fault detection equipment, pausing the failed node and cloning its operating parameters to the backup node, the problem of node failure affecting training and fine-tuning efficiency is solved, and efficient fault recovery and task continuity are achieved.

CN120704803APending Publication Date: 2025-09-26NEW H3C TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510821325.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Node failures in distributed computing clusters can affect the stability and efficiency of large model training and fine-tuning processes. Existing technologies require manual intervention and repeated data processing, resulting in low efficiency.

Method used

A model training node fault recovery method is provided. The fault detection device automatically detects node failures, pauses the failed node and clones its operating parameters to the backup node, achieving seamless switching, avoiding data loss, and ensuring task continuity and efficient execution.

Benefits of technology

It achieves automatic recovery of node failures in distributed computing clusters without manual intervention, shortens task execution time, improves fault recovery efficiency, and avoids duplication of work and data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704803A_ABST
    Figure CN120704803A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training node fault recovery method, fault detection equipment, a device and a medium, relates to the technical field of artificial intelligence, and is applied to fault detection equipment which is in communication connection with nodes in a distributed computing cluster, and the method comprises the following steps: receiving fault information sent by a first node, the fault information is sent after the first node determines that the first node has a fault and pauses execution of the target task; a data cloning instruction containing the standby node identifier is sent to the first node, so that the first node sends the operation parameters of the first node to a second node indicated by the standby node identifier, and the second node is a standby idle node; after it is detected that the second node is in the normal state, an operation starting instruction is sent to the second node, so that the second node executes the target task based on the received operation parameters. By applying the scheme provided by the embodiment of the invention, fault recovery can be carried out after the nodes in the distributed computing cluster fail.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model training node fault recovery method, fault detection equipment, device and medium. Background Art

[0002] In recent years, with the rapid development of artificial intelligence and machine learning technologies, large-scale deep learning models have been widely used in various fields. These large models typically require extensive computing power and storage resources to train and run. Due to the large number of parameters in large models, a single node cannot meet the memory requirements of large models. Therefore, large models often need to be run on distributed computing clusters. Distributed computing clusters are composed of multiple nodes, each of which is a computing device such as a server. Using distributed computing clusters, large models can be trained and fine-tuned in a relatively short time.

[0003] However, if a node in a distributed computing cluster fails, it will affect the normal execution of large model training and fine-tuning processes, thereby affecting the stability and efficiency of the training and fine-tuning processes. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a model training node fault recovery method, fault detection equipment, apparatus, and medium to perform fault recovery after a node failure in a distributed computing cluster. The specific technical solution is as follows:

[0005] In a first aspect, an embodiment of the present application provides a model training node fault recovery method, which is applied to a fault detection device, wherein the fault detection device is communicatively connected to a node in a distributed computing cluster, and the method includes:

[0006] Receiving fault information sent by the first node, wherein the fault information is sent by the first node after the first node determines that it has a fault and suspends execution of a target task, and the target task is a model training task or a model parameter adjustment task in which the first node participates;

[0007] Sending a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to a second node indicated by the standby node identifier, where the second node is a standby idle node;

[0008] After detecting that the second node is in a normal state, an operation start instruction is sent to the second node, so that the second node performs the target task based on the received operation parameters.

[0009] In one embodiment of the present application, when the first node and the third node jointly participate in executing the target task, after receiving the fault information sent by the first node, the method further includes:

[0010] Sending an operation pause instruction to the third node, so that the third node suspends execution of the target task;

[0011] After detecting that the second nodes are all in a normal state, sending an operation start instruction to the second nodes so that the second nodes perform the target task based on the received operation parameters includes:

[0012] After detecting that the second node and the third node are both in normal status, an operation start instruction is sent to the second node and the third node, so that the second node performs the target task together with the third node based on the received operation parameters.

[0013] In one embodiment of the present application, before sending the data cloning instruction including the standby node identifier to the first node, the method further includes:

[0014] Detecting whether the first node is in a fault state within a preset time period;

[0015] If the first node is not detected to be in a fault state within a preset time period, sending an operation start instruction to the first node so that the first node continues to perform the target task;

[0016] If it is detected that the first node is in a fault state within the preset time period, the step of sending a data cloning instruction including a backup node identifier to the first node is performed.

[0017] In one embodiment of the present application, the nodes in the distributed computing cluster are connected to each other via a training network and a management network, and the fault detection device is connected to the nodes in the distributed computing cluster via a management network.

[0018] The receiving fault information sent by the first node includes:

[0019] receiving fault information sent by the first node through the management network;

[0020] The sending a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier, includes:

[0021] Sending a data cloning instruction including a standby node identifier to the first node through the management network, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier through the management network;

[0022] After detecting that the second node is in a normal state, sending an operation start instruction to the second node so that the second node performs the target task based on the received operation parameters includes:

[0023] After detecting that the second node is in a normal state, an operation start instruction is sent to the second node through the management network, so that the second node interacts with other nodes participating in executing the target task through the training network based on the received operation parameters to execute the target task.

[0024] In one embodiment of the present application, the operating parameters include at least one of the following data: model parameters of the model targeted by the target task, data cached by the first node, the network address of the first node, and port configuration information of the first node, so that the second node configures itself based on the network address and the port configuration information, and continues to execute the target task based on the model parameters and the data cached by the first node.

[0025] In one embodiment of the present application, the fault information indicates that at least one of the following faults occurs in the first node: a network failure for transmitting model parameters, a network failure for accessing the storage system, a storage system failure connected to the first node, a processor failure of the first node, and the first node is unable to provide services in the distributed computing cluster.

[0026] In a second aspect, an embodiment of the present application provides a fault detection device, the fault detection device being communicatively connected to a node in a distributed computing cluster, the fault detection device comprising:

[0027] processor;

[0028] transceiver;

[0029] A machine-readable storage medium storing machine-executable instructions capable of being executed by the processor, the machine-executable instructions causing the processor to perform the following steps:

[0030] Receiving fault information sent by the first node, wherein the fault information is sent by the first node after the first node determines that it has a fault and suspends execution of a target task, and the target task is a model training task or a model parameter adjustment task in which the first node participates;

[0031] Sending a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to a second node indicated by the standby node identifier, where the second node is a standby idle node;

[0032] After detecting that the second node is in a normal state, an operation start instruction is sent to the second node, so that the second node performs the target task based on the received operation parameters.

[0033] In one embodiment of the present application, the machine executable instructions further cause the processor to perform the following steps:

[0034] Sending an operation pause instruction to the third node, so that the third node suspends execution of the target task;

[0035] After detecting that the second nodes are all in a normal state, sending an operation start instruction to the second nodes so that the second nodes perform the target task based on the received operation parameters includes:

[0036] After detecting that the second node and the third node are both in normal status, an operation start instruction is sent to the second node and the third node, so that the second node performs the target task together with the third node based on the received operation parameters.

[0037] In one embodiment of the present application, the machine executable instructions further cause the processor to perform the following steps:

[0038] Detecting whether the first node is in a fault state within a preset time period;

[0039] If the first node is not detected to be in a fault state within a preset time period, sending an operation start instruction to the first node so that the first node continues to perform the target task;

[0040] If it is detected that the first node is in a fault state within the preset time period, the step of sending a data cloning instruction including a backup node identifier to the first node is performed.

[0041] In one embodiment of the present application, the nodes in the distributed computing cluster are connected to each other via a training network and a management network, and the fault detection device is connected to the nodes in the distributed computing cluster via a management network.

[0042] The receiving the fault information sent by the first node specifically includes:

[0043] receiving fault information sent by the first node through the management network;

[0044] The sending a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier, specifically includes:

[0045] Sending a data cloning instruction including a standby node identifier to the first node through the management network, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier through the management network;

[0046] After detecting that the second node is in a normal state, sending an operation start instruction to the second node so that the second node performs the target task based on the received operation parameters specifically includes:

[0047] After detecting that the second node is in a normal state, an operation start instruction is sent to the second node through the management network, so that the second node interacts with other nodes participating in executing the target task through the training network based on the received operation parameters to execute the target task.

[0048] In one embodiment of the present application, the operating parameters include at least one of the following data: model parameters of the model targeted by the target task, data cached by the first node, the network address of the first node, and port configuration information of the first node, so that the second node configures itself based on the network address and the port configuration information, and continues to execute the target task based on the model parameters and the data cached by the first node.

[0049] In one embodiment of the present application, the fault information indicates that at least one of the following faults occurs in the first node: a network failure for transmitting model parameters, a network failure for accessing the storage system, a storage system failure connected to the first node, a processor failure of the first node, and the first node is unable to provide services in the distributed computing cluster.

[0050] In a third aspect, an embodiment of the present application provides a model training node fault recovery device, which is applied to a fault detection device, wherein the fault detection device is communicatively connected to a node in a distributed computing cluster, and the device includes:

[0051] an information receiving module, configured to receive fault information sent by a first node, wherein the fault information is sent by the first node after the first node determines that it has failed and suspends execution of a target task, wherein the target task is a model training task or a model parameter adjustment task in which the first node participates;

[0052] a cloning instruction sending module, configured to send a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to a second node indicated by the standby node identifier, where the second node is a standby idle node;

[0053] The first startup instruction sending module is used to send an operation startup instruction to the second node after detecting that the second node is in a normal state, so that the second node performs the target task based on the received operation parameters.

[0054] In one embodiment of the present application, the device further comprises:

[0055] a pause instruction sending module, configured to send an operation pause instruction to the third node, so that the third node pauses execution of the target task;

[0056] The first startup instruction sending module is specifically configured to:

[0057] After detecting that the second node and the third node are both in normal status, an operation start instruction is sent to the second node and the third node, so that the second node performs the target task together with the third node based on the received operation parameters.

[0058] In one embodiment of the present application, the device further comprises:

[0059] a fault state detection module, configured to detect whether the first node is in a fault state within a preset time period;

[0060] a second startup instruction sending module, configured to send an operation startup instruction to the first node if the first node is not detected to be in a fault state within a preset time period, so that the first node continues to perform the target task;

[0061] If it is detected that the first node is in a fault state within the preset time period, the clone instruction sending module is triggered to execute.

[0062] In one embodiment of the present application, the nodes in the distributed computing cluster are connected to each other via a training network and a management network, and the fault detection device is connected to the nodes in the distributed computing cluster via a management network.

[0063] The information receiving module is specifically used to:

[0064] receiving fault information sent by the first node through the management network;

[0065] The clone instruction sending module is specifically used to:

[0066] Sending a data cloning instruction including a standby node identifier to the first node through the management network, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier through the management network;

[0067] The first startup instruction sending module is specifically configured to:

[0068] After detecting that the second node is in a normal state, an operation start instruction is sent to the second node through the management network, so that the second node interacts with other nodes participating in executing the target task through the training network based on the received operation parameters to execute the target task.

[0069] In one embodiment of the present application, the operating parameters include at least one of the following data: model parameters of the model targeted by the target task, data cached by the first node, the network address of the first node, and port configuration information of the first node, so that the second node configures itself based on the network address and the port configuration information, and continues to execute the target task based on the model parameters and the data cached by the first node.

[0070] In one embodiment of the present application, the fault information indicates that at least one of the following faults occurs in the first node: a network failure for transmitting model parameters, a network failure for accessing the storage system, a storage system failure connected to the first node, a processor failure of the first node, and the first node is unable to provide services in the distributed computing cluster.

[0071] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the method steps described in the first aspect is implemented.

[0072] In a fifth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the methods described in the first aspect above.

[0073] Beneficial effects of the embodiments of the present application:

[0074] The model training node fault recovery method provided by the embodiment of the present application, the fault detection device based on the solution provided by the embodiment of the present application can directly control the first node to suspend operation after receiving the fault information sent by the first node, and control the first node to clone data to the idle second node, so that the second node can replace the first node after completing the data cloning and continue to complete the target task originally completed by the first node, thereby realizing the fault recovery of the node in the distributed computing cluster. This process does not require human participation and is highly efficient. In addition, after the first node determines that it has failed, it will suspend processing the target task instead of directly entering the suspended state, so the operating parameters generated before the first node fails will not be lost. Therefore, the second node can obtain the operating parameters of the first node, directly replace the first node to continue executing the target task from the breakpoint at the time of the failure, and will not return to the state before the first node fails due to the loss of operating parameters. There is no need to repeat the work that the first node has already performed before the failure, thereby shortening the time required to execute the target task. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0076] Figure 1 A flowchart of the first model training node failure recovery method provided in an embodiment of the present application;

[0077] Figure 2 A flowchart of a second model training node failure recovery method provided in an embodiment of the present application;

[0078] Figure 3 A flowchart of a third model training node failure recovery method provided in an embodiment of the present application;

[0079] Figure 4 A flowchart of a fourth model training node failure recovery method provided in an embodiment of the present application;

[0080] Figure 5 A flowchart of a fifth model training node failure recovery method provided in an embodiment of the present application;

[0081] Figure 6 A schematic diagram of the structure of a fault detection device provided in an embodiment of the present application;

[0082] Figure 7 A schematic diagram of the structure of a model training node fault recovery device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0083] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.

[0084] If a node in a distributed computing cluster fails, it will affect the normal execution of large model training and fine-tuning processes, thereby affecting the stability and efficiency of the training and fine-tuning processes. To address the above problems, embodiments of the present application provide a model training node fault recovery method, fault detection equipment, apparatus, and medium.

[0085] See also Figure 1 , which is a flow chart of the first model training node fault recovery method provided in an embodiment of the present application, and is applied to a fault detection device, and the above-mentioned fault detection device can be a server or other device with data processing capabilities. The above-mentioned fault detection device is communicatively connected to the nodes in the distributed computing cluster, and the distributed computing cluster contains multiple nodes, and each node is communicatively connected, and some or all of the nodes can jointly complete the task of model training or model parameter adjustment. Model training refers to building and optimizing a model from scratch, and through long-term training on a large amount of sample data, the model has general data processing capabilities. Model parameter adjustment refers to further short-term targeted training using sample data from a specific field on the basis of the trained model, so that the model has better data processing capabilities in the specific field. The above-mentioned node can be a server or other node with computing capabilities. The above-mentioned method includes the following steps S101-step S103.

[0086] S101: Receive fault information sent by a first node.

[0087] Among them, the above-mentioned fault information is sent after the first node determines that it has a fault and suspends the execution of the target task. The above-mentioned target task is a model training task or a model parameter adjustment task in which the above-mentioned first node participates in execution.

[0088] The first node has the ability to detect whether it has a fault. If it determines that it has a fault, it sends fault information to the fault detection device and begins to suspend the execution of the target task, but does not terminate it, so that the operating parameters of the first node can be temporarily retained. Specifically, the first node can set itself to an inactive state, thereby entering a waiting period and not processing the target task. The duration of the waiting period is a preset duration, which can be 1 hour, etc. The specific duration can be determined based on the time required for the node to complete data cloning. The preset duration should be greater than the time required to complete data cloning and achieve node replacement in most cases.

[0089] The first node may call an API (Application Programming Interface) to access the fault detection device and send fault information to the fault detection device.

[0090] Specifically, the above-mentioned fault information indicates that at least one of the following faults occurs in the above-mentioned first node: failure of the network port used to transmit model parameters, failure of the network port used to access the storage system, failure of the storage system connected to the above-mentioned first node, failure of the processor of the above-mentioned first node, and the above-mentioned first node is unable to provide services in the above-mentioned distributed computing cluster.

[0091] Specifically, the first node in the distributed computing cluster has the following status, and whether the first node fails can be determined based on the following status:

[0092] 1. Node_health_status (node ​​health status): Indicates whether the first node can provide services in the distributed computing cluster, including unhealthy and healthy states. If it is in unhealthy state, it means that the first node cannot provide services in the above distributed computing cluster.

[0093] The Node_health_status of the first node can determine whether it is in an unhealthy state or a healthy state in the following way: if the first node does not send a heartbeat signal to the distributed computing cluster manager as expected, it is determined that it is in an unhealthy state. If the resources of the first node are abnormal, the first node is in an unhealthy state. The resource abnormality refers to the utilization of the CPU (Central Processing Unit), memory or disk of the first node exceeding the threshold (for example, the threshold is 90%). If the core services such as the container, process or thread of the first node crash and fail to provide the preset function (for example, the HTTP (Hypertext Transfer Protocol) interface returns an error code), the first node is in an unhealthy state. The above examples are only some possible situations that may cause the Node_health_status of the first node to be in an unhealthy state. In the actual processing process, no matter what the reason is that the Node_health_status of the first node is in an unhealthy state, the first node sends fault information to the fault detection device.

[0094] 2. Network_status (network status): Indicates whether the network of the first node is normal, including the status of the training network and storage network. The training network is the network used by nodes in the distributed computing cluster to transmit model parameters when jointly performing model training or model parameter adjustment; the storage network is the network connecting the nodes in the distributed computing cluster and the storage system. The storage system is used to store data generated by the nodes. The storage system can be GPFS (General Parallel File System) or HDFS (Hadoop Distributed File System), etc.

[0095] If a port belonging to the training network on the first node fails or the training network is unreachable, the training network is faulty; if a port belonging to the storage network on the first node fails or the training network is unreachable, the storage network is faulty. If neither the training network nor the storage network fails, Network_status is in the healthy state; if at least one of the training network and the storage network fails, Network_status is in the error state. Therefore, if the Network_status of the first node is in the error state, the network used by the first node to access the storage system is faulty and / or the storage system connected to the first node is faulty.

[0096] 3. FS_status (File System_status, file system status): represents the status of the storage system connected to the first node.

[0097] If the storage system connected to the first node fails, the FS_status value is unhealthy. If the storage system connected to the first node is not faulty, the FS_status value is healthy. Therefore, if the first node detects that the FS_status value is unhealthy, it determines that the storage system connected to the first node has failed. If the first node is connected to GPFS, the FS_status value is GPFS_status.

[0098] 4. GPU_status (Graphics Processing Unit_status): If the processor in the first node is a GPU (Graphics Processing Unit), the first node can determine whether the first node has failed based on GPU_status. Correspondingly, if the processor in the first node is a CPU, the status indicating whether the processor in the first node has failed is CPU_status. If GPU_status is active, the GPU of the first node has not failed; if GPU_status is inactive, the GPU of the first node has failed.

[0099] S102: Send a data cloning instruction including a backup node identifier to the first node, so that the first node sends the operating parameters of the first node to the second node indicated by the backup node identifier.

[0100] The second node is a spare idle node.

[0101] The above-mentioned operating parameters include at least one of the following data: model parameters of the model targeted by the above-mentioned target task, data cached by the above-mentioned first node, the network address of the above-mentioned first node, and port configuration information of the above-mentioned first node, so that the above-mentioned second node configures itself based on the above-mentioned network address and the above-mentioned port configuration information, and continues to execute the above-mentioned target task based on the above-mentioned model parameters and the data cached by the above-mentioned first node.

[0102] After receiving the operation information, the second node can obtain the data generated by the first node during operation, thereby adjusting its own state to the state when the first node fails, and then it can replace the first node, continue to execute the target task of the first node, and complete the node replacement.

[0103] S103: After detecting that the second node is in a normal state, sending an operation start instruction to the second node, so that the second node performs the target task based on the received operation parameters.

[0104] The fault detection device can set a timeout timer for timing. The timing duration of the timeout timer is the duration of a preset period. Each time the timing reaches the timing duration, the fault detection device is triggered to detect the status of the second node. If the second node is detected to be in a normal state, it means that the second node has completed data cloning and started normal operation, and the second node has not failed and can continue to execute the target task instead of the first node. In this case, the fault detection device sends an operation start instruction to the second node to control the second node to start continuing to execute the target task.

[0105] As can be seen from the above, the fault detection device based on the solution provided by the embodiment of the present application can directly control the first node to suspend operation after receiving the fault information sent by the first node, and control the first node to clone data to the idle second node, so that the second node can replace the first node after completing the data cloning and continue to complete the target task originally completed by the first node, thereby realizing fault recovery of nodes in the distributed computing cluster. This process does not require human participation and is highly efficient. In addition, after the first node determines that it has failed, it will suspend processing the target task instead of directly entering the suspended state, so the operating parameters generated before the failure of the first node will not be lost. Therefore, the second node can obtain the operating parameters of the first node, and directly replace the first node to continue executing the target task from the breakpoint at the time of the failure, without returning to the state before the failure of the first node due to the loss of operating parameters, and there is no need to repeat the work that the first node has already performed before the failure, thereby shortening the time required to execute the target task.

[0106] Since multiple nodes are often required to jointly process tasks when using a distributed computing cluster, the first node often needs to work with other nodes to process the target task. If the first node fails, the other nodes will not be able to continue to execute the target task. Therefore, in the case where the first node and the third node jointly participate in executing the target task, see Figure 2 , is a flow chart of the second model training node failure recovery method provided in the embodiment of the present application, which is similar to the aforementioned Figure 1 Compared with the embodiment shown, the following step S104 is included after step S101, and step S103 can be implemented by the following step S103A.

[0107] S104: Sending an operation pause instruction to the third node, so that the third node pauses the execution of the target task.

[0108] The above-mentioned third node can be one or more. After a failure occurs in the first node, the fault detection device will send an operation suspension instruction to all third nodes that execute the target task together with the first node. After receiving the operation suspension instruction, the third node will set its own status to inactive and suspend processing the target task.

[0109] S103A: After detecting that the second node and the third node are both in normal status, an operation start instruction is sent to the second node and the third node, so that the second node performs the target task together with the third node based on the received operation parameters.

[0110] The fault detection device can set a timeout timer for timing. The timing duration of the timeout timer is the duration of a preset period. Each time the timing reaches the timing duration, the fault detection device is triggered to detect the status of the second node and the third node. If it is detected that the second node and the third node are both in a normal state, it means that the second node has completed data cloning and started normal operation, and neither the second node nor the third node has any faults and can continue to execute the target task. In this case, the fault detection device sends an operation start instruction to the second node and the third node to control the second node to start continuing to execute the target task.

[0111] As can be seen from the above, the embodiment of the present application takes into account that the first node in a distributed computing cluster often executes the target task together with the third node. In this case, the failure of the first node will cause the third node to be unable to continue to execute the target task. Therefore, the fault detection device will also control the third node to enter a state of suspending the execution of the target task to prevent the third node from entering a faulty state after failing to continue to execute the target task. And after the data of the first node is cloned to the second node, it is determined that both the second node and the third node are not faulty, and then the second node and the third node are controlled to continue to execute the target task, thereby ensuring the normal execution of the target task when multiple nodes participate in the execution of the target task.

[0112] See also Figure 3 , is a flow chart of the third model training node failure recovery method provided in the embodiment of the present application, which is similar to the aforementioned Figure 1 Compared with the embodiment shown, before step S102, the following step S105 is also included.

[0113] S105: Detecting whether the first node is in a fault state within a preset time period.

[0114] The fact that the fault detection device can receive the fault information sent by the first node indicates that the first node has indeed failed, but the fault may be a short-term fault, such as a communication failure between the first node and other nodes caused by short-term network fluctuations. However, the first node may have its own internally configured fault handling mechanism that can solve the fault. For example, the fault is caused by the excessively high memory usage of the first node. The first node can automatically clear the data in the memory that is not related to the target task to quickly recover from the fault; or the fault caused by environmental problems will naturally recover after the environmental conditions improve. Or the fault information reported by the first node is a false alarm. In this case, if steps S102-S103 are directly executed, although the fault handling can be completed, it takes a long time, often half an hour or even longer. However, in fact, the fault may recover by itself in a short time, or in fact, the first node does not really fail. In this case, directly cloning data will affect the execution efficiency of the target task.

[0115] Therefore, after receiving the fault information, the fault detection device will again detect whether the first node is in a faulty state, thereby determining whether the fault of the first node has been eliminated. If the first node is not detected as being in a faulty state within the preset time period, it means that the fault of the first node has been eliminated, and step S106 is executed. Conversely, if the first node is detected as being in a faulty state within the preset time period, it means that the fault of the first node has not been eliminated and is not a fault that can be eliminated in a short time, so steps S102-S103 are executed.

[0116] S106: Sending an operation start instruction to the first node, so that the first node continues to execute the target task.

[0117] Since the first node is not detected in a faulty state within the preset time, it means that the fault of the first node has been resolved, or the first node has not actually failed. Therefore, the fault detection device sends an operation start instruction to the first node, and the first node continues to execute the target task.

[0118] In another embodiment, if the first node and the third node jointly execute the target task, then after executing step S104, the fault detection device detects whether the first node and the third node are in a faulty state within a preset time period. If any of the first node and the third node are in a faulty state, the node that actually failed is controlled to clone data to an idle node. If both the first node and the third node are in a normal state within the preset time period, an operation start instruction is sent to both the first node and the third node.

[0119] As can be seen from the above, in the embodiment of the present application, after receiving the fault information sent by the first node, the fault detection device will not directly trigger the data cloning of the first node, but will re-detect whether the first node is actually in a faulty state. If the first node is still in a faulty state, data cloning will be performed again. This can avoid wasting a lot of time by directly cloning data when the fault on the first node can actually be eliminated in a short time or when the first node falsely reports fault information.

[0120] In another embodiment of the present application, the nodes in the distributed computing cluster are connected to each other through a training network and a management network, and the fault detection device and the nodes in the distributed computing cluster are connected to each other through a management network.

[0121] See also Figure 4 , is a flow chart of the fourth model training node failure recovery method provided in the embodiment of the present application, which is similar to the aforementioned Figure 1 Compared with the embodiment shown, the above step S101 can be implemented by the following step S101A, the above step S102 can be implemented by the following step S102A, and the above step S103 can be implemented by the following step S103B.

[0122] S101A: Receive fault information sent by the first node through the management network.

[0123] S102A: Sending a data cloning instruction including a backup node identifier to the first node via the management network, so that the first node sends the operating parameters of the first node to the second node indicated by the backup node identifier via the management network.

[0124] S103B: After detecting that the second node is in a normal state, an operation start instruction is sent to the second node through the management network, so that the second node can perform the target task through the training network based on the received operation parameters and interact with other nodes participating in the execution of the target task.

[0125] Among them, the other nodes participating in the target task are nodes in the distributed computing cluster that jointly complete the model training task or model parameter adjustment task indicated by the target task with the second node.

[0126] As can be seen above, the management network is used to transmit fault information and data cloning instructions between the fault detection device and the first node, and is also used to send operating parameters between the first and second nodes. The training network is used to exchange data between nodes when performing tasks. In other words, the network used during fault recovery is separated from the network used during task execution, thereby preventing the two processes from occupying network resources. More importantly, the failure of the first node may be caused by a failure in the training network. Therefore, the existence of the management network can ensure that the first node can smoothly send operating parameters to the second node in the event of a training network failure.

[0127] See also Figure 5 , which is a flowchart of the fifth model training node failure recovery method provided by the embodiment of the present application. The figure includes a model training / model parameter adjustment service running on the first node for executing the target task.

[0128] Model training / model parameter adjustment service execution: 1. Pause execution and enter a waiting state after a fault occurs. 2. Notify the fault detection device of the fault information.

[0129] The fault detection device performs 1. After receiving the fault information, it sets all nodes that execute the target task to an inactive state, and regularly checks whether the nodes set to the inactive state are faulty.

[0130] A node is inactive when it is not executing tasks.

[0131] 2. If a fault is detected, the node replacement process begins.

[0132] 3. Clone the data of the failed first node to the second node in the backup node pool.

[0133] 4. Shut down the failed first node, and after the second node is online and ready, determine that the second node is in a normal state, then set all nodes that are set to inactive state to active state.

[0134] A node is active when it continues to execute tasks.

[0135] The resource pool includes the model training / parameter adjustment pool, which includes nodes 1 and 2. The standby node pool includes standby nodes 1 and 2.

[0136] In addition, there are also solutions for fault recovery of nodes in distributed computing clusters in the related art, but fault recovery in the related art requires the user to determine whether the node has failed, manually control the node to clone data, and manually control the replaced node to continue task execution. In addition, in the related art, the node will save a checkpoint only after running a preset number of data processing times, that is, save the node's operating parameters once. However, if a node fails, it will enter a suspended state, the node's data will be lost, and subsequent fault recovery can only be based on the checkpoint. For example, a node will store a checkpoint every 100 data processing operations, but if a node fails when performing the 150th data processing, after the fault is recovered, the replaced node can only continue to execute the task based on the checkpoint from the 100th time, repeatedly wasting the node's time and computing power. However, the present application can automatically complete fault recovery without manual processing, thereby improving the efficiency of fault recovery. Furthermore, after the first node in the present application fails, it will not directly enter the suspension state, but will enter the state of suspending task execution. Therefore, the operating parameters will not be lost, and the operating parameters can be transmitted to the second node. The second node can seamlessly replace the first node to continue to execute the target task, thereby improving the efficiency of target task execution.

[0137] Corresponding to the aforementioned model training node fault recovery method, an embodiment of the present application also provides a fault detection device.

[0138] See also Figure 6 , is a schematic diagram of the structure of a fault detection device provided in an embodiment of the present application, wherein the fault detection device is communicatively connected to a node in a distributed computing cluster, and the fault detection device includes:

[0139] Processor 601;

[0140] transceiver 604;

[0141] A machine-readable storage medium 602 stores machine-executable instructions that can be executed by the processor 601, and the machine-executable instructions prompt the processor 601 to perform the following steps:

[0142] Receiving fault information sent by the first node, wherein the fault information is sent by the first node after the first node determines that it has a fault and suspends execution of a target task, and the target task is a model training task or a model parameter adjustment task in which the first node participates;

[0143] Sending a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to a second node indicated by the standby node identifier, where the second node is a standby idle node;

[0144] After detecting that the second node is in a normal state, an operation start instruction is sent to the second node, so that the second node performs the target task based on the received operation parameters.

[0145] like Figure 6 As shown, the network device may further include a communication bus 603. The processor 601, the machine-readable storage medium 602, and the transceiver 604 communicate with each other via the communication bus 603. The communication bus 603 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus 603 may be divided into an address bus, a data bus, a control bus, and the like.

[0146] The transceiver 604 may be a wireless communication module. Under the control of the processor 601 , the transceiver 604 exchanges data with other devices.

[0147] The machine-readable storage medium 602 may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Alternatively, the machine-readable storage medium 602 may be at least one storage device located remote from the processor.

[0148] The processor 601 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0149] As can be seen from the above, the fault detection device based on the solution provided by the embodiment of the present application can directly control the first node to suspend operation after receiving the fault information sent by the first node, and control the first node to clone data to the idle second node, so that the second node can replace the first node after completing the data cloning and continue to complete the target task originally completed by the first node, thereby realizing fault recovery of nodes in the distributed computing cluster. This process does not require human participation and is highly efficient. In addition, after the first node determines that it has failed, it will suspend processing the target task instead of directly entering the suspended state, so the operating parameters generated before the failure of the first node will not be lost. Therefore, the second node can obtain the operating parameters of the first node, and directly replace the first node to continue executing the target task from the breakpoint at the time of the failure, without returning to the state before the failure of the first node due to the loss of operating parameters, and there is no need to repeat the work that the first node has already performed before the failure, thereby shortening the time required to execute the target task.

[0150] In one embodiment of the present application, the machine executable instructions further cause the processor 601 to perform the following steps:

[0151] Sending an operation pause instruction to the third node, so that the third node suspends execution of the target task;

[0152] After detecting that the second nodes are all in a normal state, sending an operation start instruction to the second nodes so that the second nodes perform the target task based on the received operation parameters includes:

[0153] After detecting that the second node and the third node are both in normal status, an operation start instruction is sent to the second node and the third node, so that the second node performs the target task together with the third node based on the received operation parameters.

[0154] As can be seen from the above, the embodiment of the present application takes into account that the first node in a distributed computing cluster often executes the target task together with the third node. In this case, the failure of the first node will cause the third node to be unable to continue to execute the target task. Therefore, the fault detection device will also control the third node to enter a state of suspending the execution of the target task to prevent the third node from entering a faulty state after failing to continue to execute the target task. And after the data of the first node is cloned to the second node, it is determined that both the second node and the third node are not faulty, and then the second node and the third node are controlled to continue to execute the target task, thereby ensuring the normal execution of the target task when multiple nodes participate in the execution of the target task.

[0155] In one embodiment of the present application, the machine executable instructions further cause the processor 601 to perform the following steps:

[0156] Detecting whether the first node is in a fault state within a preset time period;

[0157] If the first node is not detected to be in a fault state within a preset time period, sending an operation start instruction to the first node so that the first node continues to perform the target task;

[0158] If it is detected that the first node is in a fault state within the preset time period, the step of sending a data cloning instruction including a backup node identifier to the first node is performed.

[0159] As can be seen from the above, in the embodiment of the present application, after receiving the fault information sent by the first node, the fault detection device will not directly trigger the data cloning of the first node, but will re-detect whether the first node is actually in a faulty state. If the first node is still in a faulty state, data cloning will be performed again. This can avoid wasting a lot of time by directly cloning data when the fault on the first node can actually be eliminated in a short time or when the first node falsely reports fault information.

[0160] In one embodiment of the present application, the nodes in the distributed computing cluster are connected to each other via a training network and a management network, and the fault detection device is connected to the nodes in the distributed computing cluster via a management network.

[0161] The receiving the fault information sent by the first node specifically includes:

[0162] receiving fault information sent by the first node through the management network;

[0163] The sending a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier, specifically includes:

[0164] Sending a data cloning instruction including a standby node identifier to the first node through the management network, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier through the management network;

[0165] After detecting that the second node is in a normal state, sending an operation start instruction to the second node so that the second node performs the target task based on the received operation parameters specifically includes:

[0166] After detecting that the second node is in a normal state, an operation start instruction is sent to the second node through the management network, so that the second node interacts with other nodes participating in executing the target task through the training network based on the received operation parameters to execute the target task.

[0167] As can be seen above, the management network is used to transmit fault information and data cloning instructions between the fault detection device and the first node, and is also used to send operating parameters between the first and second nodes. The training network is used to exchange data between nodes when performing tasks. In other words, the network used during fault recovery is separated from the network used during task execution, thereby preventing the two processes from occupying network resources. More importantly, the failure of the first node may be caused by a failure in the training network. Therefore, the existence of the management network can ensure that the first node can smoothly send operating parameters to the second node in the event of a training network failure.

[0168] In one embodiment of the present application, the operating parameters include at least one of the following data: model parameters of the model targeted by the target task, data cached by the first node, the network address of the first node, and port configuration information of the first node, so that the second node configures itself based on the network address and the port configuration information, and continues to execute the target task based on the model parameters and the data cached by the first node.

[0169] In one embodiment of the present application, the fault information indicates that at least one of the following faults occurs in the first node: a network failure for transmitting model parameters, a network failure for accessing the storage system, a storage system failure connected to the first node, a processor failure of the first node, and the first node is unable to provide services in the distributed computing cluster.

[0170] Corresponding to the aforementioned model training node fault recovery method, an embodiment of the present application also provides a model training node fault recovery device.

[0171] See also Figure 7 , is a schematic structural diagram of a model training node fault recovery device provided in an embodiment of the present application, which is applied to a fault detection device, wherein the fault detection device is communicatively connected to a node in a distributed computing cluster, and the device includes:

[0172] An information receiving module 701 is configured to receive fault information sent by a first node, wherein the fault information is sent by the first node after the first node determines that it has failed and suspends execution of a target task, wherein the target task is a model training task or a model parameter adjustment task in which the first node participates;

[0173] a cloning instruction sending module 702, configured to send a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to a second node indicated by the standby node identifier, where the second node is a standby idle node;

[0174] The first startup instruction sending module 703 is used to send an operation startup instruction to the second node after detecting that the second node is in a normal state, so that the second node performs the target task based on the received operation parameters.

[0175] As can be seen from the above, the fault detection device based on the solution provided by the embodiment of the present application can directly control the first node to suspend operation after receiving the fault information sent by the first node, and control the first node to clone data to the idle second node, so that the second node can replace the first node after completing the data cloning and continue to complete the target task originally completed by the first node, thereby realizing fault recovery of nodes in the distributed computing cluster. This process does not require human participation and is highly efficient. In addition, after the first node determines that it has failed, it will suspend processing the target task instead of directly entering the suspended state, so the operating parameters generated before the failure of the first node will not be lost. Therefore, the second node can obtain the operating parameters of the first node, and directly replace the first node to continue executing the target task from the breakpoint at the time of the failure, without returning to the state before the failure of the first node due to the loss of operating parameters, and there is no need to repeat the work that the first node has already performed before the failure, thereby shortening the time required to execute the target task.

[0176] In one embodiment of the present application, the device further comprises:

[0177] a pause instruction sending module, configured to send an operation pause instruction to the third node, so that the third node pauses execution of the target task;

[0178] The first startup instruction sending module 703 is specifically configured to:

[0179] After detecting that the second node and the third node are both in normal status, an operation start instruction is sent to the second node and the third node, so that the second node performs the target task together with the third node based on the received operation parameters.

[0180] As can be seen from the above, the embodiment of the present application takes into account that the first node in a distributed computing cluster often executes the target task together with the third node. In this case, the failure of the first node will cause the third node to be unable to continue to execute the target task. Therefore, the fault detection device will also control the third node to enter a state of suspending the execution of the target task to prevent the third node from entering a faulty state after failing to continue to execute the target task. And after the data of the first node is cloned to the second node, it is determined that both the second node and the third node are not faulty, and then the second node and the third node are controlled to continue to execute the target task, thereby ensuring the normal execution of the target task when multiple nodes participate in the execution of the target task.

[0181] In one embodiment of the present application, the device further comprises:

[0182] a fault state detection module, configured to detect whether the first node is in a fault state within a preset time period;

[0183] a second startup instruction sending module, configured to send an operation startup instruction to the first node if the first node is not detected to be in a fault state within a preset time period, so that the first node continues to perform the target task;

[0184] If it is detected that the first node is in a fault state within the preset time period, the clone instruction sending module is triggered to execute.

[0185] As can be seen from the above, in the embodiment of the present application, after receiving the fault information sent by the first node, the fault detection device will not directly trigger the data cloning of the first node, but will re-detect whether the first node is actually in a faulty state. If the first node is still in a faulty state, data cloning will be performed again. This can avoid wasting a lot of time by directly cloning data when the fault on the first node can actually be eliminated in a short time or when the first node falsely reports fault information.

[0186] In one embodiment of the present application, the nodes in the distributed computing cluster are connected to each other via a training network and a management network, and the fault detection device is connected to the nodes in the distributed computing cluster via a management network.

[0187] The information receiving module 701 is specifically configured to:

[0188] receiving fault information sent by the first node through the management network;

[0189] The clone instruction sending module 702 is specifically configured to:

[0190] Sending a data cloning instruction including a standby node identifier to the first node through the management network, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier through the management network;

[0191] The first startup instruction sending module 703 is specifically configured to:

[0192] After detecting that the second node is in a normal state, an operation start instruction is sent to the second node through the management network, so that the second node interacts with other nodes participating in executing the target task through the training network based on the received operation parameters to execute the target task.

[0193] As can be seen above, the management network is used to transmit fault information and data cloning instructions between the fault detection device and the first node, and is also used to send operating parameters between the first and second nodes. The training network is used to exchange data between nodes when performing tasks. In other words, the network used during fault recovery is separated from the network used during task execution, thereby preventing the two processes from occupying network resources. More importantly, the failure of the first node may be caused by a failure in the training network. Therefore, the existence of the management network can ensure that the first node can smoothly send operating parameters to the second node in the event of a training network failure.

[0194] In one embodiment of the present application, the operating parameters include at least one of the following data: model parameters of the model targeted by the target task, data cached by the first node, the network address of the first node, and port configuration information of the first node, so that the second node configures itself based on the network address and the port configuration information, and continues to execute the target task based on the model parameters and the data cached by the first node.

[0195] In one embodiment of the present application, the fault information indicates that at least one of the following faults occurs in the first node: a network failure for transmitting model parameters, a network failure for accessing the storage system, a storage system failure connected to the first node, a processor failure of the first node, and the first node is unable to provide services in the distributed computing cluster.

[0196] In another embodiment provided in the present application, a computer-readable storage medium is also provided, which stores a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned model training node failure recovery methods are implemented.

[0197] In another embodiment provided by the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any one of the model training node failure recovery methods in the above embodiments.

[0198] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0199] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0200] Each embodiment in this specification is described in a related manner. Similar portions between embodiments can be referenced to each other. Each embodiment focuses on the differences between other embodiments. In particular, the descriptions of the fault detection device, apparatus, machine-readable storage medium, and computer program product embodiments are relatively simple because they are generally similar to the method embodiments. For related portions, reference can be made to the descriptions of the method embodiments.

[0201] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included in the scope of protection of the present application.

Claims

1. A model training node failure recovery method, characterized in that: Applied to a fault detection device, the fault detection device being communicatively connected to a node in a distributed computing cluster, the method comprising: Receiving fault information sent by the first node, wherein the fault information is sent by the first node after the first node determines that it has a fault and suspends execution of a target task, and the target task is a model training task or a model parameter adjustment task in which the first node participates; Sending a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to a second node indicated by the standby node identifier, where the second node is a standby idle node; After detecting that the second node is in a normal state, an operation start instruction is sent to the second node, so that the second node performs the target task based on the received operation parameters.

2. The method according to claim 1, characterized in that In a case where the first node and the third node jointly participate in executing the target task, after receiving the fault information sent by the first node, the method further includes: Sending an operation pause instruction to the third node, so that the third node suspends execution of the target task; After detecting that the second nodes are all in a normal state, sending an operation start instruction to the second nodes so that the second nodes perform the target task based on the received operation parameters includes: After detecting that the second node and the third node are both in normal status, an operation start instruction is sent to the second node and the third node, so that the second node performs the target task together with the third node based on the received operation parameters.

3. The method according to claim 1, characterized in that Before sending the data cloning instruction including the standby node identifier to the first node, the method further includes: Detecting whether the first node is in a fault state within a preset time period; If the first node is not detected to be in a fault state within a preset time period, sending an operation start instruction to the first node so that the first node continues to perform the target task; If it is detected that the first node is in a fault state within the preset time period, the step of sending a data cloning instruction including a backup node identifier to the first node is performed.

4. The method according to claim 1, wherein The nodes in the distributed computing cluster are connected to each other via a training network and a management network, and the fault detection device is connected to the nodes in the distributed computing cluster via a management network. The receiving fault information sent by the first node includes: receiving fault information sent by the first node through the management network; The sending a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier, includes: Sending a data cloning instruction including a standby node identifier to the first node through the management network, so that the first node sends the operating parameters of the first node to the second node indicated by the standby node identifier through the management network; After detecting that the second node is in a normal state, sending an operation start instruction to the second node so that the second node performs the target task based on the received operation parameters includes: After detecting that the second node is in a normal state, an operation start instruction is sent to the second node through the management network, so that the second node interacts with other nodes participating in executing the target task through the training network based on the received operation parameters to execute the target task.

5. The method according to claim 1, wherein The operating parameters include at least one of the following data: model parameters of the model targeted by the target task, data cached by the first node, the network address of the first node, and port configuration information of the first node, so that the second node configures itself based on the network address and the port configuration information, and continues to execute the target task based on the model parameters and the data cached by the first node.

6. The method according to any one of claims 1 to 5, characterized in that The fault information indicates that at least one of the following faults occurs on the first node: a network failure for transmitting model parameters, a network failure for accessing a storage system, a storage system failure to which the first node is connected, a processor failure of the first node, and the first node is unable to provide services in the distributed computing cluster.

7. A fault detection device, characterized in that: The fault detection device is communicatively connected to a node in the distributed computing cluster, and the fault detection device includes: processor; transceiver; A machine-readable storage medium storing machine-executable instructions capable of being executed by the processor, the machine-executable instructions causing the processor to perform the following steps: Receiving fault information sent by the first node, wherein the fault information is sent by the first node after the first node determines that it has a fault and suspends execution of a target task, and the target task is a model training task or a model parameter adjustment task in which the first node participates; Sending a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to a second node indicated by the standby node identifier, where the second node is a standby idle node; After detecting that the second node is in a normal state, an operation start instruction is sent to the second node, so that the second node performs the target task based on the received operation parameters.

8. The fault detection device according to claim 7, characterized in that: The machine executable instructions further cause the processor to perform the following steps: Sending an operation pause instruction to the third node, so that the third node suspends execution of the target task; After detecting that the second nodes are all in a normal state, sending an operation start instruction to the second nodes so that the second nodes perform the target task based on the received operation parameters includes: After detecting that the second node and the third node are both in normal status, an operation start instruction is sent to the second node and the third node, so that the second node performs the target task together with the third node based on the received operation parameters.

9. A model training node fault recovery device, characterized in that: Applied to a fault detection device, the fault detection device being communicatively connected to a node in a distributed computing cluster, the apparatus comprising: an information receiving module, configured to receive fault information sent by a first node, wherein the fault information is sent by the first node after the first node determines that it has failed and suspends execution of a target task, wherein the target task is a model training task or a model parameter adjustment task in which the first node participates; a cloning instruction sending module, configured to send a data cloning instruction including a standby node identifier to the first node, so that the first node sends the operating parameters of the first node to a second node indicated by the standby node identifier, where the second node is a standby idle node; The first startup instruction sending module is used to send an operation startup instruction to the second node after detecting that the second node is in a normal state, so that the second node performs the target task based on the received operation parameters.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps of any one of claims 1 to 6 are implemented.