A fault handling method and related device

CN119226048BActive Publication Date: 2026-09-29HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311281352.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-06-28
Filing Date
2023-09-28
Publication Date
2026-09-29
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

[0006]目前,业界主流的故障处理方案需要较长时间恢复训练任务,效率较低,难以满足业务需求

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119226048B_ABST
    Figure CN119226048B_ABST
Patent Text Reader

Abstract

The application provides a fault processing method applied to a training system, the training system comprising a first chip on a host side and a plurality of second chips on a device side, the first chip and the plurality of second chips being used for cooperatively executing a training task, the training task comprising a first subtask and a plurality of second subtasks, execution of the second subtasks depending on an execution result of the first subtask, the method comprising: the first chip executing the first subtask, when a fault occurs, the first chip saving a fault file before a second chip on the device side in a normal state stops executing the second subtasks, and the first chip synchronizing the fault file to a chip rescheduled on the device side, so that the rescheduled chip continues to execute the second subtasks. In the method, the first chip on the host side can not stop executing the second subtasks. In this way, the execution result of the first subtask on the host side can be reused to recover the training task, the time for recovering the training task is shortened, and the efficiency of recovering the training task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202310784685.2, filed on June 28, 2023, entitled "A Fault Recovery Method for Distributed Model Training", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of artificial intelligence (AI) technology, and more particularly to a fault handling method, apparatus, scheduler, chip, computer-readable storage medium, and computer program product. Background Technology

[0003] With the continuous development of AI technology, more and more industries and fields are adopting AI models (sometimes simply referred to as models for ease of description) to achieve intelligent and automated business operations. For example, in the e-commerce industry, more and more merchants are using AI customer service built on AI models to replace human customer service, providing pre-sales and after-sales consultation services. Another example is in social networks, where platforms are using AI models to replace manual review of user-posted content, saving labor costs.

[0004] AI models are mathematical models built using AI technology to predict unknown data. For example, AI models can be object detection models or image classification models built on neural networks. AI models typically require training with large amounts of data. To improve the training efficiency of AI models, distributed training methods have emerged.

[0005] Distributed training methods distribute the training task across multiple training nodes, allowing these nodes to train the model in parallel. Executing the training task involves using a dataset to train the model and obtain its parameters (e.g., weights). Multiple training nodes can use a synchronous update mechanism to update the model's parameters. This mechanism involves accumulating the gradients obtained by each training node, calculating the average, and updating the model's parameters based on this average. When individual training nodes, the training algorithm, or the network fails, the entire distributed training task will be interrupted. As the number of training nodes increases, the likelihood of interruption also increases; therefore, a fault handling mechanism is needed to resume the training task.

[0006] Currently, mainstream fault handling solutions in the industry require a long time to recover training tasks, which is inefficient and difficult to meet business needs. Summary of the Invention

[0007] This application provides a fault handling method. This method allows the second chip on the device side to stop executing the second subtask when a fault occurs, while the first chip on the host side can continue executing the second subtask. This allows the execution result of the first subtask on the host side to be reused to resume the training task, shortening the time required to resume the training task and improving its efficiency. This application also provides an apparatus, scheduler, chip, training system, computer-readable storage medium, and computer program product corresponding to the above processing method.

[0008] Firstly, this application provides a fault handling method. This method is applied to a training system. The training system includes a first chip on the host side and multiple second chips on the device side, the first chip and the multiple second chips being used to collaboratively execute training tasks. Each training task includes a first subtask and multiple second subtasks, the execution of which depends on the execution result of the first subtask.

[0009] Specifically, the first chip executes the first subtask. When a fault occurs, the first chip saves a fault file before the second chip, which is in a normal state on the device side, stops executing the second subtask. Then, the first chip synchronizes the fault file with the rescheduled chip on the device side so that the rescheduled chip can continue executing the second subtask.

[0010] In this method, when a fault occurs, the second chip on the device side can stop executing the second subtask, while the first chip on the host side can continue executing the first subtask. This allows the execution results of the first subtask on the host side (e.g., compilation results, including but not limited to the compilation results of the computation graph) to be reused, resuming the training task, shortening the time required to resume training, and improving the efficiency of resuming training.

[0011] In some possible implementations, the fault includes recoverable chip faults or unrecoverable node faults. This method supports fine-grained recovery schemes based on different fault types, which can maximize the efficiency of resuming training tasks.

[0012] In some possible implementations, the fault is a recoverable chip failure, and the chip that the device prioritizes scheduling includes a second chip that returns to normal after a reset. This allows training tasks to resume on the original node without needing to request additional node resources, thus improving resource utilization.

[0013] In some possible implementations, the training system includes training nodes, each comprising a first chip and multiple second chips; that is, the training system can be a single-machine, multi-GPU architecture. The fault is a target chip failure among the multiple second chips. Based on this, when the first chip synchronizes the fault file with the chip prioritized for scheduling by the device, it can also synchronize the fault file with the target chip that has recovered to normal after a reset. This achieves the resumption of the training task on the original node when a recoverable chip failure occurs in a single training node.

[0014] In some possible implementations, the failure is an unrecoverable node failure, and the chip that the device prioritizes for scheduling includes a newly added third chip. This allows the training task to resume on the new node, minimizing the interruption time and improving the efficiency of training task recovery.

[0015] In some possible implementations, the training system includes a first training node and a second training node, each training node including at least one first chip and at least one second chip. That is, the training system can be a multi-machine, multi-GPU architecture. In the event of an unrecoverable node failure in the first training node, the third chip is assigned to the newly added third training node. Accordingly, the first chip of the first training node can store the execution result of the first subtask, and the first chip of the third training node can load the execution result of the first subtask.

[0016] This method caches the execution result of the first subtask and then loads the cached execution result into the new node, thereby reusing the execution result and improving the efficiency of resuming the training task.

[0017] In some possible implementations, the first subtask can be a compilation subtask (e.g., compiling the computation graph). The first chip can continue executing the first subtask, thus avoiding compilation; alternatively, the first chip of the first training node can cache the compilation results, and the first chip of the newly added third training node can load the cached compilation results from the first chip of the first training node, thus reusing the compilation results. This shortens the time required to resume the training task and improves its efficiency.

[0018] In some possible implementations, the first chip can retrieve the fault file from memory and synchronize it with the chip prioritized for scheduling by the device via a cluster communication interface. This eliminates the need for re-deserialization, shortens synchronization time, and improves fault recovery efficiency.

[0019] In some possible implementations, subtasks are executed as processes or threads. The fault handling method of this application allows processes or threads on the device side to exit, while processes or threads on the host side do not need to exit. The host side does not need to re-request resources to repeatedly execute the same subtask, further shortening the time for resuming the training task and improving the efficiency of resuming the training task.

[0020] Secondly, this application provides a fault handling method. The method is applied to a scheduler, which handles faults when a training system malfunctions. The training system includes a first chip on the host side and multiple second chips on the device side. The first chip and the multiple second chips are used to collaboratively execute a training task. The training task includes a first subtask and multiple second subtasks, the execution of which depends on the execution result of the first subtask. The method includes:

[0021] Fault detection is performed on the training system;

[0022] When a fault is detected, the first chip is notified to save a fault file before the second chip, which is in a normal state on the device side, stops executing the second subtask. The fault file is used by the first chip to synchronize to the chip that is rescheduled on the device side.

[0023] When the scheduler detects a fault, this method instructs the first chip to save the fault file before the second chip, which is functioning normally on the device side, stops executing the second subtask. The second chip on the device side can then stop executing the second subtask, while the first chip on the host side can continue executing it. This allows the execution results of the first subtask on the host side to be reused to resume the training task, shortening the recovery time and improving the efficiency of training task recovery.

[0024] In some possible implementations, the fault is a recoverable chip fault, and the scheduler can also reset the faulty chip on the device side. The chip prioritized for scheduling on the device side includes a second chip whose state returns to normal after the reset. This method supports fine-grained recovery schemes based on different fault types. For recoverable chip faults, online hot reset can be performed for recovery, and the training state can be restored by synchronizing the fault file with a normal chip.

[0025] In some possible implementations, the training system includes training nodes, each comprising a first chip and multiple second chips. A fault is identified as a target chip failure among the second chips, and a fault file is used to synchronize with the target chip via the first chip until its state returns to normal after a reset. This allows the training task to resume on the original node when a recoverable chip failure occurs at a single training node.

[0026] In some possible implementations, the failure is an unrecoverable node failure, and the chips prioritized for scheduling include the newly added third chip. Accordingly, the scheduler can also synchronize information about the newly added third chip to the normally functioning second chips. This facilitates the establishment of links between second chips, thereby aiding in the recovery of the training task.

[0027] In some possible implementations, the training system includes a first training node and a second training node, each training node including at least one first chip and at least one second chip. The first training node experiences the aforementioned unrecoverable node failure. Accordingly, the scheduler can synchronize information of the third chip to at least one second chip in the second training node. The information of the third chip includes at least one of the following: address information of the third training node to which the third chip belongs, and resource configuration information.

[0028] This method synchronizes the address, resource configuration, and other information of the third chip with at least one second chip in the second training node, enabling the third chip to resume execution of the second subtask as soon as possible and improving the efficiency of training task recovery.

[0029] In some possible implementations, the subtasks are executed as processes or threads. This method allows processes or threads on the device side to exit without exiting on the host side. The host side does not need to re-allocate resources to repeatedly execute the same subtasks, further shortening the time for resuming training tasks and improving the efficiency of training task resumption.

[0030] Thirdly, this application provides a fault handling device. The device is deployed on a first chip on the host side of a training system. The first chip and multiple second chips on the device side of the training system are used to collaboratively execute a training task. The training task includes a first subtask and multiple second subtasks, the execution of which depends on the execution result of the first subtask. The device includes:

[0031] The task execution module is used to execute the first subtask;

[0032] The file saving module is used to save the fault file when a fault occurs, before the second chip, which is in normal condition on the device side, stops executing the second subtask.

[0033] The file synchronization module is used to synchronize the faulty file with the chip that is being rescheduled by the device, so that the rescheduled chip can continue to execute the second subtask.

[0034] In some possible implementations, the fault includes recoverable chip faults or non-recoverable node faults.

[0035] In some possible implementations, the fault is a recoverable chip fault, and the chip that the device prioritizes scheduling includes a second chip whose state returns to normal after a reset.

[0036] In some possible implementations, the training system includes a training node, the training node including a first chip and a plurality of second chips, and the fault is a target chip fault among the plurality of second chips;

[0037] The file synchronization module is specifically used for:

[0038] Synchronize the fault file with the target chip whose state has returned to normal after reset.

[0039] In some possible implementations, the fault is an unrecoverable node fault, and the chip that the device prioritizes scheduling includes a newly added third chip.

[0040] In some possible implementations, the training system includes a first training node and a second training node, each training node including at least one first chip and at least one second chip. If the first training node experiences the unrecoverable node failure, the third chip belongs to a newly added third training node. The device also includes an execution result storage module for the first chip deployed in the first training node and an execution result loading module for the first chip deployed in the third training node.

[0041] The execution result storage module is used to store the execution result of the first subtask;

[0042] The execution result loading module is used to load the execution result of the first subtask.

[0043] In some possible implementations, the file synchronization module is specifically used for:

[0044] The fault file is retrieved from memory and synchronized with the chip prioritized for scheduling by the device via the cluster communication interface.

[0045] In some possible implementations, the subtasks are executed as processes or threads.

[0046] Fourthly, this application provides a scheduler. The scheduler is used for fault handling when a training system malfunctions. The training system includes a first chip on the host side and multiple second chips on the device side. The first chip and the multiple second chips are used to collaboratively execute a training task. The training task includes a first subtask and multiple second subtasks, the execution of which depends on the execution result of the first subtask. The scheduler includes:

[0047] A fault detection module is used to detect faults in the training system;

[0048] The notification module is used to notify the first chip to save a fault file before the second chip, which is in a normal state on the device side, stops executing the second subtask when a fault is detected. The fault file is used to be synchronized by the first chip to the chip that is rescheduled on the device side.

[0049] In some possible implementations, the fault is a recoverable chip fault, and the scheduler further includes:

[0050] A reset module is used to reset a faulty chip on the device side;

[0051] The chip that the device focuses on scheduling includes a second chip whose state returns to normal after a reset.

[0052] In some possible implementations, the training system includes a training node, which includes a first chip and a plurality of second chips. The fault is a target chip fault among the plurality of second chips. The fault file is used to be synchronized by the first chip to the target chip whose state has been restored to normal after a reset.

[0053] In some possible implementations, the fault is an unrecoverable node fault, the device's scheduling-focused chip includes a newly added third chip, and the scheduler further includes:

[0054] The synchronization module is used to synchronize the information of the newly added third chip to the second chip, which is in a normal state.

[0055] In some possible implementations, the training system includes a first training node and a second training node, each training node including at least one first chip and at least one second chip, wherein the first training node experiences the unrecoverable node failure.

[0056] The synchronization module is specifically used for:

[0057] The information of the third chip is synchronized to at least one second chip in the second training node. The information of the third chip includes at least one of the address information and resource configuration information of the third training node to which the third chip belongs.

[0058] In some possible implementations, the subtasks are executed as processes or threads.

[0059] Fifthly, this application provides a chip. The chip includes a processor and a memory. The processor and the memory communicate with each other. The processor executes instructions stored in the memory to cause the chip to perform the fault handling method as described in the first aspect or any implementation thereof.

[0060] Sixthly, this application provides a scheduler. The scheduler includes a processor and a memory, the memory storing computer-readable instructions; the processor executes the computer-readable instructions to cause the scheduler to perform the fault handling method as described in the first aspect or any implementation thereof.

[0061] In a seventh aspect, this application provides a training system. The training system includes a first chip on a host side and a plurality of second chips on a device side. The first chip is used to execute computer-readable instructions to perform the fault handling method as described in the first aspect or any implementation thereof.

[0062] Eighthly, this application provides a computer-readable storage medium storing instructions that instruct a computing device or a cluster of computing devices to execute the fault handling method described in the first aspect or any implementation thereof.

[0063] Ninthly, this application provides a computer program product containing instructions that, when run on a computing device or a cluster of computing devices, causes the computing device or cluster of computing devices to perform the fault handling method described in the first aspect or any implementation thereof.

[0064] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0065] To more clearly illustrate the technical methods of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below.

[0066] Figure 1A This application provides a schematic diagram of the architecture of a training system.

[0067] Figure 1B This is a schematic diagram of the architecture of another training system provided in an embodiment of this application;

[0068] Figure 2A A schematic diagram illustrating a fault handling scheme provided in an embodiment of this application;

[0069] Figure 2B A schematic diagram illustrating another fault handling solution provided in an embodiment of this application;

[0070] Figure 3 A hardware structure diagram of a server provided in an embodiment of this application;

[0071] Figure 4 A framework diagram of server deployment software provided in an embodiment of this application;

[0072] Figure 5 A flowchart illustrating a fault handling method provided in an embodiment of this application;

[0073] Figure 6 A flowchart illustrating a fault handling method provided in an embodiment of this application;

[0074] Figure 7 A schematic diagram illustrating the recovery of a training task from the original node, provided as an embodiment of this application;

[0075] Figure 8 A schematic diagram illustrating the recovery of a training task from a new node, provided as an embodiment of this application;

[0076] Figure 9 This is a schematic diagram of the structure of a fault handling device provided in an embodiment of this application;

[0077] Figure 10 This is a schematic diagram of another fault handling device provided in an embodiment of this application. Detailed Implementation

[0078] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.

[0079] First, some technical terms involved in the embodiments of this application will be introduced.

[0080] Artificial intelligence (AI), also known as machine intelligence, specifically refers to the intelligence exhibited by machines (such as computers) by imitating human thinking and behavior (such as learning, reasoning, thinking, and planning). AI typically mimics human thinking and behavior based on knowledge to achieve specific goals or complete specific tasks. This knowledge can come from experience or data.

[0081] Deep learning (DL), a branch of AI, specifically uses deep neural network models (also called deep learning models, sometimes simply referred to as models) to process massive amounts of data in order to learn knowledge (such as multiple nonlinear transformations) from the data, and then uses this knowledge to process and analyze the data. Trained deep learning models can be applied to perception and decision-making scenarios in the AI ​​field, such as image recognition, speech recognition, natural language translation, and computer gaming.

[0082] Deep learning models have a high number of parameters, typically reaching hundreds of billions or even trillions. For example, large models in the field of natural language processing (NLP) can have hundreds of billions of parameters. These large-scale deep learning models usually require massive datasets for training. A typical training method is distributed training.

[0083] Distributed training can be performed by a training system. The training system includes a first chip on the host side and multiple second chips on the device side. The first chip can be a central processing unit (CPU), also known as a central processing unit. The host side can include not only the CPU but also host memory. The second chips can be external processing units used to assist the CPU in completing tasks and accelerate computation. Therefore, the second chip can also be called an accelerator card or accelerator. The second chip can include a neural processing unit (NPU) or a graphics processing unit (GPU). Different types of processors / processing units can implement different types of instruction set architectures. Similar to the host side, the device side can also include device memory. In some examples, the device memory can be integrated with the second chips on the device side.

[0084] The training system can employ a single-machine multi-GPU architecture or a multi-machine multi-GPU architecture to achieve distributed training. It should be noted that in the above x-machine x-GPU configuration, "machine" refers to the host machine, and "GPU" refers to devices such as accelerator cards. The system architecture of the training system is described below with reference to the accompanying diagram.

[0085] See Figure 1A , Figure 1B The diagram shown illustrates the architecture of the training system. Figure 1A In this system, the training system can be a single training node, which includes a host 102 and multiple devices 104. The host 102 includes a first chip 1022 and host memory 1024, and the devices 104 can be second chips. Figure 1B The training system may include multiple training nodes 100, such as a first training node and a second training node. Each training node 100 includes a host 102 and at least one device 104. The host 102 includes a first chip 1022 and host memory 1024, and the device 104 may be a second chip.

[0086] In this setup, the first chip 1022 can be a CPU, and the second chip can be an NPU or a GPU. The first chip 1022 and the second chip can be connected via a bus. The first chip 1022 and multiple second chips can work together to complete the training task. For example, the training task can be broken down into multiple sub-tasks, with different chips executing different sub-tasks. Based on this, the aforementioned training task is also called a distributed training task.

[0087] It should be noted that, Figure 1A , Figure 1BThe system architecture shown is merely exemplary, and other architectures may be used in other possible implementations of this application embodiment. For example, host 102 may also include multiple CPUs, which can be connected to form a mesh.

[0088] The training system can use an iterative method to update the model parameters to achieve model training. Each iteration updates the model parameters once; an iteration can also be called a training step, or simply a step. The number of samples used in each iteration is called the batch size. During training, the process of using all the sample data in the dataset (e.g., the training set) once is called an epoch. For ease of understanding, the following example illustrates this. In this example, the training set includes 1000 sample data points, and the batch size can be 100. Therefore, each iteration uses 100 sample data points, and 10 iterations with the 1000 sample data points in the training set complete one epoch of training.

[0089] In distributed training mechanisms, training nodes (or chips used for training) can update model parameters synchronously. This synchronous update mechanism involves accumulating the gradients obtained from each training node (or chip) to calculate the mean, and then updating the model parameters based on this mean. Compared to asynchronous update mechanisms, where each training node updates the model parameters based on its own gradient, synchronous update mechanisms ensure a more stable decrease in loss, avoiding significant fluctuations.

[0090] Synchronous update mechanisms bind computation and communication to each gradient synchronization. However, with synchronous update mechanisms, the entire distributed training task will be interrupted if an individual training node, training algorithm, or network fails. The likelihood of interruption increases with the number of training nodes. To address this, related technologies provide solutions for backing up faulty files to recover the training task.

[0091] like Figure 2A As shown, a typical fault handling scheme involves an elastic agent detecting faults. Upon detecting a fault, the agent notifies the corresponding processes in the same process group to stop. The elastic agent then reinitializes the cluster. The training task creates a new ProcessGroup on the non-faulty cluster nodes, loads periodic fault files (e.g., checkpoint (CKPT) files), and continues training based on these CKPT files. This prevents catastrophic failures caused by server maintenance or network problems, ensuring no loss of training progress.

[0092] like Figure 2B As shown, another typical fault handling scheme is for the scheduling component to perform fault detection. The types of faults detected include chip faults and node faults (including parameter plane faults, such as cluster communication faults). When a fault is detected, the node where the fault occurred is isolated, the training process is stopped, and the AI ​​framework saves a terminal CKPT file (the CKPT file at the time of the fault, also known as the breakpoint CKPT file). The training task can be rescheduled to a non-faulty cluster node, the terminal CKPT file is loaded, and the training job can continue.

[0093] However, the above scheme will terminate all processes upon detecting a fault, regardless of the type of fault (also known as stopping, specifically by closing the file descriptors opened by the processes to release the resources they occupy). This means that the training system needs to restart processes to perform corresponding tasks, such as recompiling, when different types of faults occur, which greatly affects the efficiency of resuming training tasks.

[0094] In view of this, this application provides a fault handling method. This method is applied to a training system, which includes a first chip on the host side and multiple second chips on the device side. The first chip and the multiple second chips are used to collaboratively execute a training task. The training task includes a first subtask and multiple second subtasks, and the execution of the second subtasks depends on the execution result of the first subtask. For example, the first subtask may be a subtask of compiling a computation graph, and the second subtasks may be subtasks of executing an operator sequence based on the compilation result of the computation graph.

[0095] Specifically, the first chip executes the first subtask. When a fault occurs, the first chip saves the fault file before the second chip, which is in normal condition on the device side, stops executing the second subtask. Then, the first chip synchronizes the fault file with the chip that is rescheduled on the device side so that the rescheduled chip can continue to execute the second subtask.

[0096] In this method, when a fault occurs, the second chip on the device side can stop executing the second subtask, while the first chip on the host side can continue executing the first subtask. This allows the execution result of the first subtask on the host side (e.g., compilation result, including but not limited to the compilation result of the computation graph) to be reused, resuming the training task, shortening the time for resuming the training task, and improving its efficiency. Furthermore, the subtask can be executed through a process or thread. In other words, the fault handling method of this application allows the process or thread on the device side to exit, while the process or thread on the host side can remain running. The host side does not need to re-allocate resources to repeatedly execute the same subtask, further shortening the time for resuming the training task and improving its efficiency.

[0097] Furthermore, this method supports fine-grained recovery schemes based on different fault types. For recoverable chip faults, online hot reset can be performed for recovery, and the training state can be restored by synchronizing the fault file with normal chips. For unrecoverable node faults, only the subtasks on the faulty node stop execution (process or thread stops) and the subtasks resume execution on the new training node. The training state is synchronized by synchronizing the fault file with normal training nodes.

[0098] The fault handling method of this application can be applied to various distributed training scenarios. For example, this fault handling method can be used in the scenario of distributed training of image recognition models, where different network structures of the image recognition model (e.g., different sub-networks or different network layers) can be trained by different training nodes or multiple second chips of a single training node. When a training node or a second chip (or some second chips) fails, the second chip on the device side can stop executing the second subtask, while the first chip on the host side can continue executing the second subtask. Specifically, the first chip saves a fault file before the normally functioning second chips stop executing the second subtask, and then synchronizes the fault file with the rescheduled chip on the device side so that the rescheduled chip can continue executing the second subtask.

[0099] This fault handling method can also be applied to scenarios involving distributed training of text recognition models, where different network structures can be trained separately by different training nodes or a second chip. The fault handling mechanism in case of a failure can be referenced from the scenario for distributed training of image recognition models, and will not be elaborated further here.

[0100] Figure 1A , Figure 1B The framework of the training system has been introduced. The following section uses a single-machine multi-GPU architecture as an example to introduce the hardware structure and software logic of the training system.

[0101] A training system with a single-machine, multi-GPU architecture can be a server, while a training system with a multi-machine, multi-GPU architecture can be a server cluster. A server cluster is a group of multiple servers, and its hardware structure can be similar to that of a single server. Specifically, users can purchase or lease servers, which can be cloud servers or physical servers.

[0102] See Figure 3The diagram illustrates a server hardware structure, where server 30 includes a host 32 and a device 34. The host 32 and device 34 are connected. The host 32 includes a processor and memory (i.e., host memory). The processor can be a CPU, and the memory can be a Dual In-line Memory Module (DIMM). Specifically, the DIMM can be of double data rate (DDR) type, such as DDR4 DIMM. Figure 3 In the example, host 32 includes 4 CPUs and 4 DDR4 DIMM groups, with each CPU connected to one DDR4 DIMM group, and each DDR4 DIMM group including 8 DDR4 DIMMs. Multiple CPUs of host 32 can be connected to form a hydra mesh.

[0103] Optionally, host 32 may also include one or more interfaces, such as a Serial Advanced Technology Attachment (SATA) interface, a Next-Generation Non-Volatile Memoryexpress (NVMe) interface, and a Gigabit Ethernet (GE) interface. Host 32 may also include memory. The memory may include SATA-enabled memory or NVMe-enabled memory, such as a SATA-enabled hard disk drive (HDD) or an NVMe-enabled solid-state drive (SSD).

[0104] Device 34 includes an accelerator card. Figure 3 In the example, device 34 may include an accelerator such as an NPU. Figure 3 The example of device 34 including 8 NPUs is used for illustration. In other possible implementations of the embodiments of this application, device 34 may also include more accelerator cards, such as more types of accelerator cards or more numbers of accelerator cards.

[0105] Then, see Figure 4The diagram shows the framework of the software deployed on the server. Users can install firmware 302 and driver 304 on server 30. Firmware 302 is typically a program written to read-only memory that can directly control and interact with the hardware, and check for any hardware errors. Driver 304 is a small piece of code added to the operating system, containing information about the hardware. When a computer program requests to interact with a piece of hardware, driver 304 can act as a translator of instructions between the hardware and the program using it. For example, firmware 302 can control and interact with device 34, and check for any errors on device 34; driver 304 can act as a translator of instructions between device 34 and the program using it.

[0106] Furthermore, when the hardware architecture of server 30 adopts a heterogeneous computing architecture (including computing architectures using processing units with different types of instruction sets), users can also install a heterogeneous computing framework 306 on server 30. In distributed training scenarios, the heterogeneous computing framework 306 can be a heterogeneous computing framework for neural networks (Compute Architecture for Neuro Net, CANN). CANN can support users in quickly building AI applications by providing multi-level programming interfaces. Here, AI applications refer to applications built based on the trained AI models. It should be noted that the heterogeneous computing framework 306 is an optional framework. The fault handling method of the embodiments of this application can still be executed even if the above framework is not installed on server 30. The role of the above framework is to improve the efficiency of building AI applications.

[0107] Then, the user can install a deep learning framework 308 on server 30. Deep learning framework 308 is used to compile the computation graph of the model and automatically perform gradient calculations within the computation graph. Thus, during distributed training, the graph compilation results (such as operator sequences) can be executed to perform related computations for distributed training. Depending on the compilation method, deep learning frameworks can be divided into frameworks that support static compilation and frameworks that support dynamic compilation. Users can choose to install one or more deep learning frameworks 308 on server 30 according to their business needs. In some embodiments, deep learning framework 308 may not be installed on server 30; in this case, server 30 can implement the model from scratch using a programming language such as Python.

[0108] Users can also install a scheduler 310 on server 30. This scheduler 310 is used to schedule training tasks to achieve distributed training. The scheduler 310 can be a distributed scheduling component (distributed scheduling framework) or an AI development platform. In some embodiments, the distributed scheduling component or AI development platform can be a self-developed component or platform, or it can be a third-party distributed scheduling component or a third-party AI development platform.

[0109] Furthermore, users can deploy a model library 312 on server 30. This model library 312 includes AI models implemented using a unified framework, which have standardized parameters and APIs. The AI ​​models include reusable configuration items defined within the unified framework. This reduces the configuration work required for the AI ​​models.

[0110] The scheduler 310 can launch training tasks on training nodes, download training code from model library 312, and read configuration files from the training code downloaded from model 312. When a training node fails or one or more second chips (such as the NPU) in the training node fail, the scheduler 310 can notify the first chip (such as the CPU) to save a fault file before the second chip with normal status on the device side stops executing the second subtask. This fault file is used by the first chip to synchronize to the chip that is rescheduled on the device side.

[0111] The scheduler 310 can be provided to the user as a code package, which the user can install or deploy themselves. Alternatively, the scheduler 310 can be provided to the user as a cloud service. Specifically, the cloud service provider can provide an application programming interface (API) for fault handling, and the training system can call the API to implement the fault handling method of this application.

[0112] To make the technical solution of this application clearer and easier to understand, the fault handling method of this application will be described below with reference to the accompanying drawings.

[0113] See Figure 5 The flowchart illustrates a fault handling method applied to a training system. The training system includes a first chip on the host side and multiple second chips on the device side. The first chip and the multiple second chips collaboratively execute a training task, which includes a first subtask and multiple second subtasks. The execution of the second subtasks depends on the execution result of the first subtask. For example, the first subtask may be a compilation subtask, and its execution result may be a compilation result. The second subtasks are operator sequence (kernel sequence) execution subtasks. The execution of the operator sequence depends on the compilation result on the host side. The method includes the following steps:

[0114] S502: The first chip executes the first subtask. If a fault occurs, S504 is executed.

[0115] The first subtask is the subtask in the training task that executes host-side logic. The training task specifically involves iteratively executing the model's computation graph. A computation graph is a high-level representation that provides a global view of operators without specifying implementation details. To maximize computational resource reuse, improve efficiency, and shorten training time, the computation graph typically requires compilation optimization. Compilation optimization may include high-level data writing to generate an optimized computation graph, followed by operator-level optimization to generate efficient code for fusion operators in the computation graph. Operators are declared using a tensor expression language. Then, an overhead model is used to search for optimal operator code from a potential optimization set, and the generated code is packaged into deployable modules, such as kernel sequences. This compilation optimization process can be executed on the host side, while kernel sequences can be executed on the device side. Therefore, the first subtask may include a compilation optimization subtask (also called a compilation subtask). The second subtask may include a subtask that executes kernel sequences.

[0116] When iterating through the computational graph of the model, it is usually necessary to read the dataset and perform data preprocessing on the data within the dataset. The data preprocessing operations can vary depending on the data type of the dataset. For example, if the dataset includes image data, data preprocessing operations may include cropping, scaling, and rotation. Or, if the dataset includes text data, data preprocessing operations may include word segmentation. This data preprocessing can be performed on the host side. Accordingly, the first subtask may also include a data preprocessing subtask.

[0117] The first chip can execute the first subtask through a process or a thread. A process or thread is an application running in memory. A process is an encapsulation of a runtime program and is the basic unit of resource scheduling and allocation. Threads typically depend on processes to achieve concurrency within a process. Processes have independent memory, while threads share the process's memory. Taking a compiler optimization subtask as an example, the first chip can execute the compiler optimization subtask through one or more processes. In some possible implementations, the first chip can also execute the compiler optimization subtask through one or more threads.

[0118] S504. Before the second chip, which is in a normal state on the device side, stops executing the second subtask, the fault file is saved.

[0119] The failures that occur in the training system can be either recoverable chip failures or unrecoverable node failures. Recoverable chip failures refer to those that can be recovered through a thermal reset. Unrecoverable node failures refer to failures that are difficult to recover from at the original node. Unrecoverable node failures can include, but are not limited to, failures of all chips on the device side of the training node or parameter plane failures. Parameter plane failures can include cluster communication failures.

[0120] Because the training system includes multiple second chips, when a second chip fails, the first chip can save the fault file before the second chip in a normal state stops executing the second subtask. In some possible implementations, for multi-machine, multi-GPU training systems, if a training node in the training system experiences a recoverable chip failure or an unrecoverable node failure, the first chip of the normal node (the training node that did not fail) can save the fault file before the second chip of the normal node stops executing the subtask. Specifically, the first chip of the normal node calls the interface of the heterogeneous computing architecture component through an AI framework (such as a deep learning framework), causing the heterogeneous computing architecture component to exit, the process or thread based on the heterogeneous computing architecture component to stop, and the second subtask to stop executing. The AI ​​framework can save the fault file before the second subtask stops executing. In other possible implementations, for single-machine, multi-GPU training systems, if a training node experiences an unrecoverable chip failure, the first chip of that training node can save the fault file before the second chip in that node that did not fail stops executing the second subtask. In some possible implementations, the first subtask may not stop executing; in other words, the process or thread used to execute the first subtask may not exit.

[0121] The first chip can read fault files from the device side and store them in the host-side memory when saving fault files, without needing to persist them to disk. In other words, the first chip does not need to write fault files to disk; for example, it does not write fault files to Object Storage Service (OBS) or the Open Computing Kit (OCK). Furthermore, the first chip can save near-terminal fault files, such as near-terminal CKPTs (also known as breakpoint CKPTs), thus preventing the loss of training progress.

[0122] S506, The first chip synchronizes the fault file with the chip that is rescheduled by the device so that the rescheduled chip can continue to execute the second sub-task.

[0123] When a training task is interrupted due to a fault, it can be rescheduled to a chip that is in normal working order; this rescheduled chip is called the rescheduled chip. The first chip synchronizes the fault file with the rescheduled chip, enabling the rescheduled chip to continue executing the second subtask, thereby resuming the training task.

[0124] For recoverable chip failures, the rescheduled chip can be a second chip that recovers to normal operation after a reset. For example, in a training system with a single-machine, multi-GPU architecture, the training node includes a first chip and multiple second chips. If the target chip among the second chips fails, the rescheduled chip can be the target chip that recovers to normal operation after a reset. Accordingly, the first chip can synchronize the fault file with the second chip (such as the target chip) that has recovered to normal operation after a reset.

[0125] For unrecoverable node failures, the rescheduled chip can be a newly added third chip. For example, a training system includes a first training node and a second training node, each including at least one first chip and at least one second chip. If the first training node experiences an unrecoverable node failure, the training task can be rescheduled to a newly added third training node, and the third chip belongs to the newly added third training node. Accordingly, the first chip can synchronize the fault file with the third chip (or the third training node).

[0126] It should be noted that the first chip of the first training node can also store the execution result of the first subtask, such as the cached compilation result. The first chip of the third training node can load the execution result of the first subtask without recompiling.

[0127] When the fault file is stored in memory (e.g., host memory), the first chip can read the fault file from memory and then synchronize the fault file with the chip prioritized for scheduling via the cluster interface. Model information (such as model weights) is no longer serialized and deserialized via CKPT, avoiding the impact of model size, input / output (IO), and bandwidth on training task recovery. Instead, point-to-point state synchronization is achieved through the cluster mechanism.

[0128] When the first chip saves a near-failure file such as a near-failure CKPT, the first chip can synchronize the training status (or training progress) at the time of the failure, thus avoiding the loss of training process caused by periodic CKPT.

[0129] Based on the above description, it can be seen that the fault handling method of this application supports the following: when a fault occurs, the second chip on the device side can stop executing the second subtask, while the first chip on the host side can continue executing the second subtask. This allows the execution result of the first subtask on the host side (e.g., compilation result, including but not limited to the compilation result of the computation graph) to be reused, resuming the training task, shortening the time for resuming the training task, and improving the efficiency of resuming the training task.

[0130] The above describes the fault handling method from the perspective of the first chip. The following describes the fault handling method of this application from the perspective of the scheduler.

[0131] See Figure 6 The flowchart illustrates a fault handling method applied to a scheduler. The scheduler handles faults in the training system. The training system includes a first chip on the host side and multiple second chips on the device side. The first chip and the multiple second chips collaboratively execute training tasks. Each training task includes a first subtask and multiple second subtasks, with the execution of the second subtasks depending on the execution result of the first subtask. The method includes the following steps:

[0132] S602, The scheduler performs fault detection on the training system.

[0133] Specifically, the scheduler can perform fault detection on the second chip or the training node where the second chip is located in the training system. Specifically, the scheduler can detect faults in the second chip through heartbeat messages. The second chip in the training system can periodically send heartbeat messages to the scheduler, and the scheduler determines whether the second chip has failed based on the reception of these heartbeat messages. For example, if the scheduler does not receive a heartbeat message from the second chip for N consecutive periods, it indicates that the second chip or the node where the second chip is located has failed. The scheduler can also determine the fault type by combining the reception of heartbeat messages from other second chips on the same training node.

[0134] For K second chips of the same training node, if the scheduler does not receive heartbeat messages from any of the K second chips of that training node, for example, if it does not receive heartbeat messages from the aforementioned K second chips for N consecutive cycles, then the scheduler can determine that the training node has failed. For the same training node, if the scheduler does not receive heartbeat messages from some of the K second chips, but receives heartbeat messages from the remaining second chips, then the scheduler can determine that the second chip for which the corresponding heartbeat message was not received has failed.

[0135] It should be noted that the heartbeat message is only one way to implement fault detection in the training system. In other possible implementations of this application, the scheduler can also implement fault detection in other ways. For example, the scheduler can detect faults in the second chip by reading its status code.

[0136] S604. When a fault is detected, the scheduler notifies the first chip to save the fault file before the second chip, which is in normal condition on the device side, stops executing the second subtask.

[0137] The fault file is used by the first chip to synchronize with the chip that is being rescheduled by the device. In this way, the rescheduled chip can continue to execute the second subtask based on the fault file synchronized by the first chip.

[0138] In some possible implementations, for recoverable chip failures, the scheduler can also reset the faulty chip on the device side, for example, through an online hot reset. Accordingly, the chips prioritized for scheduling on the device side include a second chip whose state returns to normal after the reset.

[0139] In other possible implementations, for unrecoverable node failures, the device-side scheduling includes a newly added third chip. The scheduler can synchronize the information of the newly added third chip to the normal second chips. The information of the third chip is used to establish links with other second chips participating in training. Specifically, the training system includes a first training node and a second training node. Each training node includes at least one first chip and at least one second chip. The first training node experiences an unrecoverable node failure. Training tasks can be rescheduled to the third training node. The third chip is the device-side chip of the third training node. The scheduler can synchronize the information of the third chip to at least one second chip in the second training node. The information of the third chip includes at least one of the following: the address information of the third training node to which the third chip belongs, and resource configuration information. The address information of the third training node can be the IP address of the third training node, or other addresses that can be used to establish links. The resource configuration information can be a sorting table of the third chips, which records the sequence number of the third chip occupied by the process or thread executing the subtask.

[0140] Based on the above description, it can be seen that the fault handling method of this application, when the scheduler detects a fault, notifies the first chip to save the fault file before the second chip, which is in normal condition on the device side, stops executing the second subtask. The second chip on the device side can stop executing the second subtask, while the first chip on the host side can continue executing the second subtask. In this way, the execution result of the first subtask on the host side can be reused to restore the training task, shortening the time for restoring the training task and improving the efficiency of restoring the training task.

[0141] The fault detection method of this application has been introduced from the perspectives of the first chip and the scheduler. The fault handling method of this application will be explained below in conjunction with specific application scenarios.

[0142] First, see Figure 7 The diagram illustrates a scenario where the original node is used to recover the training task. Figure 7 As shown, the training system includes multiple training nodes, which include a first chip on the host side and a second chip (such as an NPU / GPU) on the device side. Figure 7 The example illustrates that each training node's device side includes 8 second chips. If a recoverable chip failure occurs in the second chip numbered 0-3 of a training node, the training task can be resumed on the original node.

[0143] In this system, training nodes that experience a malfunction are called abnormal nodes, and nodes that do not malfunction are called normal nodes. Abnormal and normal nodes perform corresponding operations to resume the training task on the original node.

[0144] For abnormal nodes, the distributed scheduling framework (such as the scheduler) deployed on the host side detects a chip failure and resets the faulty chip (e.g., the second chip with a rank of 0-3). The chip failure or reset can trigger the exit of heterogeneous computing architecture components on the device side; correspondingly, the process on the device side can exit, and the second subtask stops executing. The corresponding process of the host-side AI framework does not exit, thus eliminating the need to recompile the computation graph, while waiting for the chip reset and the exit of the heterogeneous computing architecture components. Once the host-side AI framework detects that the second chip has returned to normal after the reset, it can trigger a link-building process to re-establish a link with the normal node, for example, by using a collection communication library. Then, the host-side AI framework re-executes the computation graph, waiting for the faulty file to be synchronized. When synchronization is complete, the abnormal node can continue training the job. For example, the second chip of the abnormal node can continue executing the second subtask.

[0145] For normal nodes, the distributed scheduling framework deployed on the host side can instruct the AI ​​framework to call the heterogeneous computing architecture component interface, causing the heterogeneous computing architecture component to exit, the process or thread based on the heterogeneous computing architecture component to stop, and the second subtask to stop execution. Before the process or thread based on the heterogeneous computing architecture component stops (the second subtask stops execution), the AI ​​framework saves the fault file. After saving the fault file, the heterogeneous computing architecture component can be quickly decommissioned. The corresponding process or thread of the AI ​​framework does not exit, thus eliminating the need for recompilation. Once the host-side AI framework detects that the chip status is normal, it can re-establish the chain. Then, the AI ​​framework re-executes the computation graph, synchronizing the fault file with the recovered nodes through the cluster communication interface according to the fault file recovery strategy. Normal nodes can then continue training jobs.

[0146] Next, see Figure 8 The diagram illustrates a scenario where a new node is used to resume the training task. Figure 8 As shown, the training system includes multiple training nodes, which include a first chip on the host side and a second chip (such as an NPU / GPU) on the device side. Figure 8 The example illustrates that each training node includes 8 secondary chips on its device side. If a training node experiences an unrecoverable node failure (all secondary chips on that training node fail), the training task can be resumed on a new node.

[0147] Abnormal nodes, normal nodes, and new nodes each perform corresponding operations to resume the training task on the new node.

[0148] For abnormal nodes, the distributed scheduling framework detects the failure, and the heterogeneous computing architecture component can exit. The AI ​​framework can cache the compilation results based on the return value (abnormal return value) of the exiting heterogeneous computing architecture component. Then, the AI ​​framework process can exit.

[0149] For healthy nodes, the distributed scheduling framework instructs the AI ​​framework to call the heterogeneous computing architecture component interface to cause the heterogeneous computing architecture component to exit. Before the heterogeneous computing architecture component exits, the AI ​​framework saves the fault file to the host side. After saving the file, the heterogeneous computing architecture component can exit quickly. Processes or threads based on the heterogeneous computing architecture component can exit. The second subtask stops executing. The host-side AI framework does not exit, so recompilation is not required. The host-side AI framework checks that the chip status is normal and can re-establish the chain. Then, the AI ​​framework re-executes the computation graph. According to the fault file recovery strategy, the AI ​​framework synchronizes the fault file to the recovered nodes through the cluster communication interface. Healthy nodes can continue training jobs.

[0150] For new nodes, the distributed scheduling framework synchronizes the new node's IP and rank table to each container group (POD) in each cluster. The AI ​​framework obtains the new node's IP and updates the rank table. Then, the AI ​​framework loads the cached compilation results from the faulty node, for example, using OCK to accelerate the loading of the cached compilation results. Next, the AI ​​framework checks that the chip status is normal and re-establishes the chain. Then, the AI ​​framework can re-execute the computation graph, waiting for the fault file synchronization to complete. After the fault file synchronization is complete, the newly added node can continue the training job.

[0151] Based on the aforementioned fault handling method, this application also provides a fault handling apparatus. For example... Figure 9As shown, the fault handling device 900 is deployed on a first chip on the host side of the training system. The first chip and multiple second chips on the device side of the training system are used to collaboratively execute training tasks. The training task includes a first subtask and multiple second subtasks. The execution of the second subtasks depends on the execution result of the first subtask. The fault handling device 900 includes:

[0152] Task execution module 902 is used to execute the first subtask;

[0153] The file saving module 904 is used to save the fault file when a fault occurs, before the second chip, which is in normal condition on the device side, stops executing the second subtask.

[0154] The file synchronization module 906 is used to synchronize the fault file with the chip that is being rescheduled by the device, so that the rescheduled chip can continue to execute the second subtask.

[0155] For example, the task execution module 902, file saving module 904, and file synchronization module 906 described above can be implemented in hardware or in software.

[0156] When implemented in software, the task execution module 902, file saving module 904, and file synchronization module 906 can be applications running on computing devices, such as computing engines. Applications can also be virtualized through virtualization services to be provided to users. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, and container services. VM services can be services that use virtualization technology to create virtual machine (VM) resource pools on multiple physical hosts to provide VMs for users to use on demand. BMS services are services that create virtual BMS resource pools on multiple physical hosts to provide BMS for users to use on demand. Container services are services that create virtual container resource pools on multiple physical hosts to provide containers for users to use on demand. A VM is a simulated virtual computer, that is, a logical computer. A BMS is a scalable, high-performance computing service with computing performance indistinguishable from traditional physical machines and features secure physical isolation. Containers are a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service, and container service mentioned above are merely specific examples. In practical applications, virtualization services can also include other lightweight or heavyweight virtualization services, which are not specifically limited here.

[0157] When implemented in hardware, the task execution module 902, file saving module 904, and file synchronization module 906 may include at least one computing device, such as a server. Alternatively, the task execution module 902, file saving module 904, and file synchronization module 906 may also be devices implemented using application-specific integrated circuits (ASICs) or programmable logic devices (PLDs). The aforementioned PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0158] In some possible implementations, the fault includes recoverable chip faults or non-recoverable node faults.

[0159] In some possible implementations, the fault is a recoverable chip fault, and the chip that the device prioritizes scheduling includes a second chip whose state returns to normal after a reset.

[0160] In some possible implementations, the training system includes a training node, the training node including a first chip and a plurality of second chips, and the fault is a target chip fault among the plurality of second chips;

[0161] The file synchronization module 906 is specifically used for:

[0162] Synchronize the fault file with the target chip whose state has returned to normal after reset.

[0163] In some possible implementations, the fault is an unrecoverable node fault, and the chip that the device prioritizes scheduling includes a newly added third chip.

[0164] In some possible implementations, the training system includes a first training node and a second training node, each training node including at least one first chip and at least one second chip. If the first training node experiences the unrecoverable node failure, the third chip belongs to a newly added third training node. The device 900 also includes an execution result storage module 908 for the first chip deployed in the first training node and an execution result loading module 909 for the first chip deployed in the third training node.

[0165] The execution result storage module 908 is used to store the execution result of the first subtask;

[0166] The execution result loading module 909 is used to load the execution result of the first subtask.

[0167] The execution result saving module 908 and the execution result loading module 909 can be implemented in software or hardware. When implemented in software, the execution result saving module 908 and the execution result loading module 909 can be applications running on a computing device, such as a computing engine. When implemented in hardware, the execution result saving module 908 and the execution result loading module 909 can include at least one computing device, such as a server. Alternatively, the execution result saving module 908 and the execution result loading module 909 can also be devices implemented in ASIC or PLD.

[0168] In some possible implementations, the file synchronization module 906 is specifically used for:

[0169] The fault file is retrieved from memory and synchronized with the chip prioritized for scheduling by the device via the cluster communication interface.

[0170] In some possible implementations, the subtasks are executed as processes or threads.

[0171] Based on the aforementioned fault handling method, this application also provides a scheduler. The scheduler is used for fault handling when a training system malfunctions. The training system includes a first chip on the host side and multiple second chips on the device side. The first chip and the multiple second chips are used to collaboratively execute training tasks. Each training task includes a first subtask and multiple second subtasks, the execution of which depends on the execution result of the first subtask. Figure 10 As shown, the scheduler 1000 includes:

[0172] Fault detection module 1002 is used to detect faults in the training system;

[0173] The notification module 1004 is used to notify the first chip to save a fault file before the second chip, which is in a normal state on the device side, stops executing the second subtask when a fault is detected. The fault file is used to be synchronized by the first chip to the chip that is rescheduled on the device side.

[0174] For example, the fault detection module 1002 and the notification module 1004 described above can be implemented in hardware or in software.

[0175] When implemented in software, the fault detection module 1002 and notification module 1004 can be applications running on a computing device. When implemented in hardware, the fault detection module 1002 and notification module 1004 can include at least one computing device, such as a server. Alternatively, the fault detection module 1002 and notification module 1004 can also be devices implemented using an ASIC or a PLD. Alternatively, the fault detection module 1002 and notification module 1004 can also be a processing unit within the aforementioned devices.

[0176] In some possible implementations, the fault is a recoverable chip fault, and the scheduler 1000 further includes:

[0177] The reset module 1006 is used to reset the faulty chip on the device side;

[0178] The chip that the device focuses on scheduling includes a second chip whose state returns to normal after a reset.

[0179] The reset module 1006 can be implemented in hardware or in software.

[0180] When implemented in software, the reset module 1006 can be an application running on a computing device. When implemented in hardware, the reset module 1006 can include at least one computing device, such as a server. Alternatively, the reset module 1006 can be an ASIC-implemented or PLD-implemented device. Alternatively, the reset module 1006 can also be a processing unit within the aforementioned devices.

[0181] In some possible implementations, the training system includes a training node, which includes a first chip and a plurality of second chips. The fault is a target chip fault among the plurality of second chips. The fault file is used to be synchronized by the first chip to the target chip whose state has been restored to normal after a reset.

[0182] In some possible implementations, the fault is an unrecoverable node fault, the device-focused scheduling chip includes a newly added third chip, and the scheduler 1000 further includes:

[0183] The synchronization module 1008 is used to synchronize the information of the newly added third chip to the second chip, which is in normal condition.

[0184] Similarly, the synchronization module 1008 can be implemented in hardware or in software.

[0185] When implemented in software, the synchronization module 1008 can be an application running on a computing device. When implemented in hardware, the synchronization module 1008 can include at least one computing device, such as a server. Alternatively, the synchronization module 1008 can be an ASIC-implemented or PLD-implemented device. Alternatively, the synchronization module 1008 can also be a processing unit within the aforementioned devices.

[0186] In some possible implementations, the training system includes a first training node and a second training node, each training node including at least one first chip and at least one second chip, wherein the first training node experiences the unrecoverable node failure.

[0187] The synchronization module 1008 is specifically used for:

[0188] The information of the third chip is synchronized to at least one second chip in the second training node. The information of the third chip includes at least one of the address information and resource configuration information of the third training node to which the third chip belongs.

[0189] This application also provides a chip. The chip may include a processor and a memory. The memory stores computer-readable instructions; the processor executes the computer-readable instructions to perform the fault handling method in the foregoing embodiments. For example, the chip may be a first chip, and the processor of the first chip executes the computer-readable instructions to perform the fault handling method executed by the first chip in the foregoing embodiments.

[0190] This application also provides a scheduler. The scheduler may include a processor and a memory. The memory stores computer-readable instructions; the processor executes the computer-readable instructions to perform the fault handling method executed by the scheduler in the foregoing embodiments, for example... Figure 6 The method of the illustrated embodiment.

[0191] This application also provides a training system. The training system may include a first chip on the host side and multiple second chips on the device side. The training system can be a single-machine, multi-GPU architecture, meaning it can be a single training node, such as a training server. This training node may include a host (containing the first chip) and multiple accelerator cards (i.e., second chips). The first chip on the host side and the multiple second chips on the device side can collaboratively execute distributed training tasks. In some examples, the training system may also be a multi-machine, multi-GPU architecture, meaning it can be a cluster of multiple training nodes. Each of the multiple training nodes includes a first chip on the host side and at least one second chip on the device side. The multiple training nodes can collaboratively execute distributed training tasks. The first chip in the training system is used to execute computer-readable instructions to perform the fault handling method of the aforementioned embodiments.

[0192] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the execution of the fault handling method described in the foregoing embodiments.

[0193] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the aforementioned fault handling method.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A fault handling method, characterized in that, The method is applied to a training system, which includes a first chip on the host side and multiple second chips on the device side. The first chip and the multiple second chips are used to collaboratively execute a training task. The training task includes a first subtask and multiple second subtasks, and the execution of the second subtasks depends on the execution result of the first subtask. The first chip executes the first sub-task; When a fault occurs, the first chip reads the fault file from the device side before the second chip, which is in normal condition on the device side, stops executing the second subtask, and saves the fault file in the host side memory; The first chip synchronizes the fault file with the chip that is rescheduled by the device, so that the rescheduled chip can continue to execute the second subtask.

2. The method according to claim 1, characterized in that, The faults include recoverable chip faults or non-recoverable node faults.

3. The method according to claim 1, characterized in that, The fault is a recoverable chip fault, and the chip that the device prioritizes scheduling includes a second chip whose state returns to normal after a reset.

4. The method according to any one of claims 1 to 3, characterized in that, The training system includes a training node, the training node includes a first chip and a plurality of second chips, and the fault is a target chip fault among the plurality of second chips; The first chip synchronizes the fault file with the chip prioritized for scheduling by the device, including: The first chip synchronizes the fault file with the target chip, whose state has returned to normal after a reset.

5. The method according to claim 1 or 2, characterized in that, The fault is an unrecoverable node fault, and the chip that the device focuses on scheduling includes the newly added third chip.

6. The method according to claim 5, characterized in that, The training system includes a first training node and a second training node. Each training node includes at least one first chip and at least one second chip. If the first training node experiences an unrecoverable node failure, the third chip belongs to a newly added third training node. The method further includes: The first chip of the first training node stores the execution result of the first subtask; The first chip of the third training node loads the execution result of the first subtask.

7. The method according to any one of claims 1 to 3, characterized in that, The first chip synchronizes the fault file with the chip prioritized for scheduling by the device, including: The first chip retrieves the fault file from the host-side memory and synchronizes the fault file with the chip scheduled on the device side via the cluster communication interface.

8. The method according to any one of claims 1 to 3, characterized in that, The first subtask and / or the second subtask are executed in a process or thread manner.

9. A fault handling method, characterized in that, The method is applied to a scheduler used for fault handling when a training system malfunctions. The training system includes a first chip on the host side and multiple second chips on the device side. The first chip and the multiple second chips are used to collaboratively execute training tasks. Each training task includes a first subtask and multiple second subtasks, the execution of which depends on the execution result of the first subtask. Fault detection is performed on the training system; When a fault is detected, the first chip is notified to read the fault file from the device side before the second chip, which is in a normal state on the device side, stops executing the second subtask. The fault file is stored in the host side memory and is used by the first chip to synchronize to the chip that is rescheduled on the device side.

10. The method according to claim 9, characterized in that, The fault is a recoverable chip fault, and the method further includes: Reset the faulty chip on the device side; The chip that the device focuses on scheduling includes a second chip whose state returns to normal after a reset.

11. The method according to claim 9, characterized in that, The training system includes a training node, which includes a first chip and multiple second chips. The fault is a target chip fault among the multiple second chips. The fault file is used to be synchronized by the first chip to the target chip whose state has been restored to normal after a reset.

12. The method according to claim 9, characterized in that, The fault is an unrecoverable node fault, the chip that the device prioritizes for scheduling includes the newly added third chip, and the method further includes: The information of the newly added third chip is synchronized to the second chip, which is in a normal state.

13. The method according to claim 12, characterized in that, The training system includes a first training node and a second training node, each training node including at least one first chip and at least one second chip, wherein the first training node experiences the unrecoverable node failure. The process of synchronizing the information of the newly added third chip to the second chip, which is in a normal state, includes: The information of the third chip is synchronized to at least one second chip in the second training node. The information of the third chip includes at least one of the address information and resource configuration information of the third training node to which the third chip belongs.

14. The method according to any one of claims 9 to 13, characterized in that, The first subtask and / or the second subtask are executed in a process or thread manner.

15. A fault handling device, characterized in that, The device is deployed on a first chip on the host side of the training system. The first chip and multiple second chips on the device side of the training system are used to collaboratively execute training tasks. The training task includes a first subtask and multiple second subtasks. The execution of the second subtasks depends on the execution result of the first subtask. The device includes: The task execution module is used to execute the first subtask; The file saving module is used to read the fault file from the device side before the second chip, which is in normal condition on the device side, stops executing the second subtask when a fault occurs, and save the fault file in the memory on the host side. The file synchronization module is used to synchronize the faulty file with the chip that is being rescheduled by the device, so that the rescheduled chip can continue to execute the second subtask.

16. The apparatus according to claim 15, characterized in that, The faults include recoverable chip faults or non-recoverable node faults.

17. The apparatus according to claim 15, characterized in that, The fault is a recoverable chip fault, and the chip that the device prioritizes scheduling includes a second chip whose state returns to normal after a reset.

18. The apparatus according to any one of claims 15 to 17, characterized in that, The training system includes a training node, the training node includes a first chip and a plurality of second chips, and the fault is a target chip fault among the plurality of second chips; The file synchronization module is specifically used for: Synchronize the fault file with the target chip whose state has returned to normal after reset.

19. The apparatus according to claim 15 or 16, characterized in that, The fault is an unrecoverable node fault, and the chip that the device focuses on scheduling includes the newly added third chip.

20. The apparatus according to claim 19, characterized in that, The training system includes a first training node and a second training node. Each training node includes at least one first chip and at least one second chip. If the first training node experiences an unrecoverable node failure, the third chip belongs to a newly added third training node. The device also includes an execution result storage module for the first chip deployed in the first training node and an execution result loading module for the first chip deployed in the third training node. The execution result storage module is used to store the execution result of the first subtask; The execution result loading module is used to load the execution result of the first subtask.

21. The apparatus according to any one of claims 15 to 17, characterized in that, The file synchronization module is specifically used for: The fault file is retrieved from the host-side memory and synchronized with the chip scheduled on the device side via the cluster communication interface.

22. The apparatus according to any one of claims 15 to 17, characterized in that, The first subtask and / or the second subtask are executed in a process or thread manner.

23. A scheduler, characterized in that, The scheduler is used for fault handling when a training system malfunctions. The training system includes a first chip on the host side and multiple second chips on the device side. The first chip and the multiple second chips are used to collaboratively execute training tasks. The training task includes a first subtask and multiple second subtasks. The execution of the second subtasks depends on the execution result of the first subtask. The scheduler includes: A fault detection module is used to detect faults in the training system; The notification module is used to notify the first chip to read the fault file from the device side before the second chip, which is in a normal state on the device side, stops executing the second subtask when a fault is detected. The fault file is stored in the memory on the host side and is used by the first chip to synchronize the fault file to the chip that is rescheduled on the device side.

24. The scheduler according to claim 23, characterized in that, The fault is a recoverable chip fault, and the scheduler further includes: A reset module is used to reset a faulty chip on the device side; The chip that the device focuses on scheduling includes a second chip whose state returns to normal after a reset.

25. The scheduler according to claim 23, characterized in that, The training system includes a training node, which includes a first chip and multiple second chips. The fault is a target chip fault among the multiple second chips. The fault file is used to be synchronized by the first chip to the target chip whose state has been restored to normal after a reset.

26. The scheduler according to claim 23, characterized in that, The fault is an unrecoverable node fault. The chip that the device focuses on scheduling includes a newly added third chip. The scheduler also includes: The synchronization module is used to synchronize the information of the newly added third chip to the second chip, which is in a normal state.

27. The scheduler according to claim 26, characterized in that, The training system includes a first training node and a second training node, each training node including at least one first chip and at least one second chip, wherein the first training node experiences the unrecoverable node failure. The synchronization module is specifically used for: The information of the third chip is synchronized to at least one second chip in the second training node. The information of the third chip includes at least one of the address information and resource configuration information of the third training node to which the third chip belongs.

28. The scheduler according to any one of claims 23 to 27, characterized in that, The first subtask and / or the second subtask are executed in a process or thread manner.

29. A chip, characterized in that, The device includes a processor and a memory, the memory storing computer-readable instructions; the processor executes the computer-readable instructions to cause the chip to perform the method as described in any one of claims 1 to 8.

30. A scheduler, characterized in that, It includes a processor and a memory, the memory storing computer-readable instructions; the processor executes the computer-readable instructions to cause the scheduler to perform the method as described in any one of claims 9 to 14.

31. A training system, characterized in that, The training system includes a first chip on the host side and a plurality of second chips on the device side, wherein the first chip is used to perform the method as described in any one of claims 1 to 8.

32. A computer program product, characterized in that, Includes computer-readable instructions; the computer-readable instructions are used to implement the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Model training method, server, chip and system

    CN114936117A

  • Fault file storage method and related device

    CN114968947A