Ai model training method, distributed training system, and related device
By storing AI model parameter values in different areas in memory and persisting storage in case of failure, the training loss problem caused by computing node failure is solved, and the model training efficiency and the availability of distributed systems are improved.
Patent Information
- Application Number
- PCT/CN2024/115835
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2024-08-30
- Publication Date
- 2025-07-31
AI Technical Summary
During the AI model training process, the training loss caused by computing node failure is longer, which reduces the overall efficiency of model training and the availability of distributed training systems.
By storing parameter values generated during two adjacent rounds of training in different storage areas in memory, and starting the fault saving mechanism when a fault occurs, the parameter values are persisted so as to quickly restore the model to the pre-failure state and reduce the training loss time.
It effectively reduces the training loss time after failure recovery, improves the overall efficiency of model training and the availability of distributed training systems, and ensures training accuracy and efficiency.
Smart Images

Figure CN2024115835_31072025_PF_FP_ABST
Abstract
Description
AI model training methods, distributed training systems, and related equipment
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 23, 2024, with application number 202410100931.2 and application name “AI model training method, distributed training system and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence technology, and in particular to an AI model training method, a distributed training system, and related equipment. Background Art
[0003] With the development of artificial intelligence (AI) technology, there are various ways to train AI models. Specifically, when the parameter scale of the AI model is small, a single computing node can be used to iteratively train the AI model. When the parameter scale of the AI model is large, an AI cluster including multiple computing nodes can be used to jointly train the AI model. For example, for an AI model with more than 100 billion parameters, a dozen or dozens of computing nodes can be used to complete the model training.
[0004] During the training of the AI model, computing nodes may fail. For example, as the scale of the AI cluster increases, the failure rate of the AI cluster also increases proportionally. For this reason, a checkpoint mechanism is usually used to train the AI model, and the AI model training is restored after a failure occurs. As shown in Figure 1, during the training of the AI model, the parameter values in the currently trained AI model are periodically backed up. In this way, if a failure occurs during the training of the AI model, the AI model can continue to be trained based on the most recently backed up parameter values (such as the parameter values obtained after the yth round of training in Figure 1), so there is no need to retrain the AI model from scratch.
[0005] However, the AI model restored based on the backup parameter values is usually separated from the AI model at the time of the failure by one or more rounds of iterative training. After the failure is recovered, it is necessary to resume training from the AI model that has undergone round y training, that is, the AI model needs to re-execute the training process between rounds y+1 and z. The time consumed by this process is the time lost during the entire model training process, hereinafter referred to as the training loss time. Therefore, the training loss time generated after the failure is recovered will reduce the overall efficiency of the model training, wherein the longer the training loss time, the lower the overall efficiency of the model training.
[0006] Summary of the Invention
[0007] This application provides an AI model training method to reduce the training loss time after fault recovery and improve the overall efficiency of model training. In addition, this application also provides a distributed training system, computing device, computer-readable storage medium, and computer program product.
[0008] In a first aspect, the present application provides an AI model training method that can be executed by a computing node, hereinafter referred to as a first computing node. Taking two rounds of iterative training of an AI model as an example, the AI model is hosted on the first computing node. During the Nth (N is a positive integer) round of training of the AI model, the first computing node stores the first state data of the AI model in a first storage area in the memory. The first state data includes the parameter values of the AI model after the Nth round of training. During the N+1th round of training of the AI model, the first computing node gradually updates the parameter values in the AI model. For example, if the AI model may include multiple network layers, the AI model can update the parameter values of the next network layer after the parameter values of one network layer are updated. Furthermore, for the parameter values that have been updated, the first computing node does not use the parameter values to overwrite the parameter values in the first storage area, but instead stores the updated parameter values in the second storage area in the memory to prevent the parameter values stored in the first storage area from being adjusted. Moreover, when a failure occurs during the N+1 round of AI model training, such as a failure in a software program involved in AI model training, the first computing node starts a fault saving mechanism, which includes persistently storing the first state data in the first storage area in the memory, such as persistently storing the first state data in a local hard disk, or persistently storing it in cloud storage, etc.
[0009] Since the first computing node stores the parameter values obtained from the training in different storage areas in the memory during the Nth round of training and the N+1th round of training of the AI model during the iterative training of the AI model, this allows, when the AI model is trained for the N+1th round, if a failure occurs during the process of updating the parameter values, the first computing node can use the parameter values saved in the first storage area to restore the AI model to the AI model after completing the Nth round of training. This can effectively reduce the training loss time after a failure occurs during the model training process (that is, the training loss time is less than the time required for one round of training), thereby improving the overall efficiency of AI model training. When the first computing node belongs to a distributed training system, this AI model training method can also effectively improve the availability of the distributed training system.
[0010] Moreover, the newly generated parameter values of two adjacent rounds of training of the AI model can be saved to different storage areas in the memory, which enables the parameter values to be updated uniformly, that is, it can avoid the situation where the values of some parameters in the AI model are updated while the values of the remaining parameters are not updated (such as failure to be updated due to a malfunction, etc.), thereby avoiding the values of the parameters in the AI model not being updated uniformly and affecting the training accuracy of the AI model, thereby ensuring that the training accuracy of the AI model can reach a high level. In addition, during the iterative training of the AI model, the parameter values obtained through different rounds of training will be stored in different storage areas in the memory, without the need to copy the parameter values in the memory, which can avoid the copying of the parameter values affecting the efficiency of the training of the AI model by the first computing node.
[0011] In one possible implementation, the fault preservation mechanism may further include, after the first computing node completes fault recovery, restoring the AI model after completing the Nth round of training based on the persistently stored first state data, and continuing to perform the N+1th round of training on the restored AI model. In this way, after the first computing node completes fault recovery, it can continue to execute the AI model training process from the N+1th round, without having to train the AI model from scratch. At the same time, the training loss time caused by the failure of the first computing node can be less than the time required for one round of training, thereby effectively improving the training efficiency of the AI model.
[0012] In one possible embodiment, the first computing node belongs to a distributed training system, which includes multiple computing nodes (and may also include at least one switching node, so that multiple computing nodes can communicate through the switching node). The multiple computing nodes in the distributed training system can be used to jointly train the AI large model, wherein each computing node can carry part of the model in the AI model, or each computing node can carry the completed AI large model. The above-mentioned first computing node is a computing node among the multiple computing nodes included in the distributed training system, wherein the AI model carried on the first computing node is part or all of the model of the AI large model. In this way, in the process of jointly training the AI large model using multiple computing nodes, each computing node can use different storage areas in the memory to store the parameter values generated during two adjacent rounds of model training, which can effectively reduce the training loss time after a failure occurs during the model training process, thereby improving the overall efficiency of AI model training and improving the availability of the distributed training system.
[0013] Optionally, the first computing node can also train the AI model independently without jointly training the AI model with other computing nodes.
[0014] In one possible embodiment, the multiple computing nodes in the distributed training system include not only a first computing node but also a second computing node. The second computing node and the first computing node are configured with the same AI model. In this case, the first computing node and the second computing node can train the AI model at the same pace, that is, the second computing node and the first computing node can simultaneously store the same state data in memory. Then, when a failure occurs during the M+1th (M is a positive integer) round of training of the AI model, the first computing node obtains second state data from the second computing node. The second state data includes the parameter values of the AI model after the Mth round of training. After the first computing node completes fault recovery, it can restore the AI model after the Mth round of training based on the second state data obtained from the second computing node and continue to perform the M+1th round of training on the AI model. In this way, in a data-parallel distributed training scenario, in addition to restoring the AI model by persistently storing the state data, the first computing node can also restore the AI model by obtaining state data from other computing nodes. This can improve the reliability of restoring the AI model after the first computing node fails and thus ensure the training efficiency of the AI model.
[0015] In one possible implementation, the multiple computing nodes in the distributed training system include not only a first computing node but also a third computing node. In this case, the AI model carried by the first computing node is a partial model of the large AI model to be trained, and the AI model carried by the third node is another partial model of the large AI model. In other words, the distributed training system can deploy different parts of the large AI model on different computing nodes through model parallelism, with each computing node responsible for training a portion of the large AI model, thereby reducing the computing power requirements of a single computing node.
[0016] In one possible implementation, the failure of the first computing node may specifically be a failure of a software program on the first computing node that participates in training the AI model, such as a software program execution error, thereby interrupting AI model training. Alternatively, the failure of the first computing node may be a communication failure sent by the first computing node. For example, in a data-parallel training scenario, the first computing node needs to exchange gradients with other computing nodes during parameter update. Therefore, if the communication function of the first computing node is abnormal, the AI model training may be interrupted.
[0017] In one possible implementation, the first state data stored in the memory of the first computing node may include, in addition to parameter values, momentum. This momentum is used to update the parameter values of the AI model during the N+1 round of AI model training. In practical applications, if the memory space of the first computing node is sufficient, storing data other than the AI model parameter values can improve the subsequent recovery and training of the AI model.
[0018] In one possible implementation, the first computing node is an accelerator card, and the memory used by the first computing node to store the first state data is HBM (High Bandwidth Memory). In actual applications, the first computing node may also be a CPU or a server, and the memory used to store the first state data may be other types of storage media.
[0019] In a second aspect, the present application provides a distributed training system, which includes multiple computing nodes, and the multiple computing nodes are used to jointly train an AI large model, wherein the multiple computing nodes include a first computing node, and the AI model carried on the first computing node is part or all of the AI large model; the first computing node is used to: during the Nth round of training of the AI model, store the first state data of the AI model in a first storage area in the memory, the first state data including the parameter value of the AI model after the Nth round of training, the AI model is carried on the first computing node, and N is a positive integer; during the N+1th round of training of the AI model, gradually update the parameter values in the AI model, and store the updated parameter values in the second storage area in the memory; when a fault occurs during the N+1th round of training of the AI model, start the fault saving mechanism, and the fault saving mechanism includes persistently storing the first state data in the first storage area in the memory.
[0020] In one possible implementation, the fault saving mechanism further includes: after the first computing node completes fault recovery, recovering the AI model after completing the Nth round of training based on the persistently stored first state data; and continuing to perform the N+1th round of training on the AI model.
[0021] In one possible embodiment, the multiple computing nodes also include a second computing node, which is configured with the same AI model as the first computing node; the first computing node is further used to: when a failure occurs during the M+1 round of training of the AI model, obtain second state data from the second computing node, the second state data including the parameter values of the AI model after the Mth round of training, where M is a positive integer; after the first computing node completes fault recovery, restore the AI model after the Mth round of training according to the second state data; and continue to perform the M+1 round of training on the AI model.
[0022] In one possible implementation, the multiple computing nodes further include a third computing node, the AI model carried by the first computing node is a partial model in the AI large model, and the AI model carried by the third computing node is another partial model in the AI large model.
[0023] In one possible implementation, the fault is a fault occurring in a software program involved in training the AI model on the first computing node, or the fault is a communication fault occurring in the first computing node.
[0024] In a possible implementation, the first state data further includes momentum, and the momentum is used to update the parameter values of the AI model during the N+1th round of training of the AI model.
[0025] In a possible implementation, the first computing node is an accelerator card, and the memory is HBM (High Bandwidth Memory).
[0026] The distributed training system provided in the second aspect corresponds to the AI model training method provided in the first aspect. Therefore, the technical effects of the second aspect and the various implementation methods in the second aspect can be referred to the technical effects of the corresponding implementation methods in the first aspect, and no further details will be given.
[0027] In a third aspect, the present application provides an accelerator card, which is used to execute the AI model training method described in the first aspect and any implementation method of the first aspect.
[0028] In a fourth aspect, the present application provides a computing device comprising a processor and a memory; the memory is used to store instructions, and the processor executes the instructions stored in the memory, so that the computing device executes the AI model training method described in the above-mentioned first aspect and any one of the implementation methods of the first aspect.
[0029] In a fifth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computing device, the computing device executes the operating steps of the AI model training method described in the first aspect and any one of the implementation methods of the first aspect.
[0030] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when run on a computing device, enables the computing device to execute the operating steps of the AI model training method described in the first aspect and any one of the implementation methods of the first aspect.
[0031] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG1 is a schematic diagram of the structure of a distributed training system;
[0033] FIG2 is a schematic diagram of the structure of an exemplary distributed training system provided by the present application;
[0034] FIG3 is a schematic diagram of the structure of another exemplary distributed training system provided by the present application;
[0035] FIG4 is a flow chart of an AI model training method provided in this application;
[0036] Figure 5 is a schematic diagram of reducing the training loss time;
[0037] FIG6 is a flow chart of an AI model training method provided in this application;
[0038] FIG7 is a schematic diagram of restoring model training using state data stored in persistent storage;
[0039] FIG8 is a flow chart of another AI model training method provided in this application;
[0040] FIG9 is a schematic diagram of resuming model training using state data stored in computing node 102;
[0041] FIG10 is a schematic diagram of the hardware structure of a computing device provided in this application. DETAILED DESCRIPTION
[0042] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, various non-limiting embodiments of the embodiments of the present application will be exemplified below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, not all of them. Based on the embodiments in this application, all other embodiments obtained based on the above content are within the scope of protection of this application.
[0043] FIG2 is a schematic diagram illustrating the structure of an exemplary distributed training system 20. As shown in FIG2 , the distributed training system 20 may include multiple computing nodes. Furthermore, the distributed training system may also include at least one switching node, and the multiple computing nodes may be communicatively connected to the at least one switching node. For ease of understanding, FIG2 illustrates an example system comprising four computing nodes (computing nodes 101 to 104) and three switching nodes (switching nodes 201 to 203).
[0044] Among them, the computing node can be a node with model training capability. Exemplarily, the computing node can be an accelerator card, which can be, for example, a deep learning processor (deep-learning processing unit, DPU), a data processing unit (Data processing unit, DPU), a graphics processing unit (graphics processing unit, GPU), a neural network processor (neural-network processing unit, NPU), or a tensor processing unit (tensor processing unit, TPU), etc., or other types of accelerator cards. Alternatively, the computing node can be a general-purpose processor, such as a central processing unit (CPU), etc. Alternatively, the computing node can also include a computing device of a CPU and an accelerator card. This application does not limit the specific implementation of the computing node.
[0045] Furthermore, each computing node is configured with memory. For example, as shown in FIG2 , computing node 101 may include memory 1012, and may also include a computing unit 1011 for providing computing power and a communication unit 1013 for communicating with other computing nodes. The remaining computing nodes may have the same structure as computing node 101, and will not be described in detail here. Exemplarily, the memory in the computing node may be, for example, high bandwidth memory (HBM), or may be other types of memory, and this is not limited.
[0046] The switching node may be a node with data forwarding capability, such as a switch.
[0047] In the distributed training system 20, multiple computing nodes can be used to perform distributed training on the AI large model. The AI large model refers to a model with a relatively large parameter scale, such as a model with a parameter amount greater than a threshold. The threshold can be set according to the needs of the actual application scenario, for example, the threshold can be 100,000. In this embodiment, there is no limitation on the parameter scale in the AI large model. In different application scenarios, whether a model is an AI large model can be judged based on thresholds of different sizes. Exemplarily, the AI model can be a neural network model, such as a natural language processing model, a generative reasoning model, etc. In actual application, the AI model can also be other types of models.
[0048] In the first implementation, the number of parameters in the large AI model meets preset conditions, such as the large AI model may include trillions of parameters. In this case, the distributed training system 20 can adopt a model parallel approach, using multiple computing nodes to jointly train the same large AI model. Different computing nodes are used to train different parts of the large AI model. That is, the AI model carried by each computing node can be a part of the large AI model. Taking the large AI model as a neural network model as an example, the network layers included in the neural network model can be divided into multiple groups, each group including at least one network layer, and each group can be scheduled to a different computing node. Accordingly, each computing node can train part of the network layer in the large AI model. In this way, the model training task on each computing node can adapt to the computing power of the computing node, thereby improving the training efficiency of the large AI model and reducing the computing power requirements of a single computing node by training different parts of the large AI model in parallel through multiple computing nodes.
[0049] In the second implementation, the number of parameters in the AI large model does not meet the preset conditions, such as the AI large model may include hundreds of thousands of parameters. In this case, the distributed training system 20 can use data parallelism to train the same AI large model in parallel using multiple computing nodes. That is, each computing node can be configured with the same AI large model and different data sets, so that multiple computing nodes can train the same AI large model using different data sets. During each round of training the AI large model, each computing node can obtain the gradients obtained by the training of the other computing nodes, aggregate the gradients obtained by the training of multiple computing nodes, and update the model parameters based on the aggregated results, thereby completing a round of model training.
[0050] In actual application, multiple computing nodes in the distributed training system 20 can also use other methods to complete distributed training for large AI models, such as combining the first implementation method and the second implementation method mentioned above, etc., and there is no limitation on this.
[0051] During the iterative training of the AI large model, the distributed training system 20 may malfunction, such as anomalies in the software programs involved in training the AI large model on some computing nodes, or interruptions in the communication connection between some computing nodes and the switching nodes. If the distributed training system 20 periodically backs up the parameter values obtained by training each computing node, not only will the periodic backup of the parameter values reduce the efficiency of the distributed training system 20 in training the AI large model, but also, when the distributed training system 20 fails, the training time lost by restoring the AI large model based on the backup parameter values will also reduce the efficiency of the distributed training system 20 in training the AI large model.
[0052] Based on this, in the distributed training system 20 shown in Figure 2, each computing node can store the parameter values obtained after different rounds of training in different storage areas in the memory during the training of the AI model. Among them, the AI model carried by each computing node can be part or all of the models in the AI large model to be trained. Specifically, taking computing node 101 as an example, during the Nth round of training of the AI model carried on it, computing node 101 can store the parameter values of the trained AI model in the storage area A in the memory 1012, where N is a positive integer. Then, computing node 101 continues to train the AI model for the N+1th round. During this process, computing node 101 can gradually update the parameter values in the AI model, and for the parameter values that have been updated, computing node 101 does not store the updated parameter values in storage area A (that is, it does not replace the parameter values before the update with the updated parameter values in storage area A), but stores the updated parameter values in storage area B in memory 1012. That is, the different values of the parameters in the AI model before and after the update are respectively stored in different storage areas in the memory 1012. In this way, when a failure occurs during the N+1 round of training of the AI model and the gradual updating of the parameter values, the computing node 101 can persistently store the parameter values stored in the storage area A so that the parameter values can be used to restore the AI model later, that is, restore the AI model to the state after the Nth round of training. In this way, the training loss time caused by failures during the model training process can be effectively reduced. Specifically, the training loss time can be controlled within a range that is less than the time required for one round of training of the AI model, such as controlling the training loss time within a few seconds or tens of seconds, thereby improving the overall efficiency of model training.
[0053] Moreover, reducing the training loss time can also effectively improve the availability of the distributed training system 20. The availability of the distributed training system 20 can be calculated by the following formula (1). 训练损失 ) / (MTBF + MTTR) Formula (1)
[0054] Among them, “Availability” refers to the availability of the distributed training system 20; “MTBF” refers to the mean time between failures (MTBF), that is, the average time interval between two adjacent failures in the distributed training system 20; “T 训练损失", refers to the average training loss time, which can be half of the maximum value of the training loss time; "MTTR", refers to the mean time to recovery (MTTR), which refers to the average time required to recover from a fault in the distributed training system 20. As shown in formula (1), T 训练损失 Therefore, by reducing the training loss time in the distributed training system 20, the availability of the distributed training system 20 can be effectively improved.
[0055] Furthermore, the computing node 101 can alternately use storage area A and storage area B in the memory 1012 to store updated parameter values generated during the iterative training process. For example, during the N+2 round of training of the AI model, the computing node 101 can write the newly generated parameter values into storage area A. That is, after the parameter values obtained after the N+1 round of training have been successfully stored in storage area B, the parameter values obtained after the N+2 round of training will be used to overwrite the parameter values obtained after the Nth round of training stored in storage area A. This process continues in this way until the computing node 101 completes the training process for the AI model.
[0056] In addition, during each round of training of the AI model, the computing node 101 can uniformly update the parameter values in the AI model, that is, it can avoid the situation where some parameter values in the AI model finally trained by the computing node 101 are updated while the remaining parameter values are not updated due to a failure. In actual application scenarios, if there are some rounds of model training in which only some parameter values in the AI model are updated and the remaining parameter values are not updated, the updates of the parameter values in the AI model will affect each other, which will cause the unupdated parameter values to have unexpected effects in subsequent iterative training processes, such as causing the AI model to fail to converge or causing the AI model to have low training accuracy. Therefore, the computing node 101 restores the AI model based on the parameter values stored in the storage area A, which can ensure that the parameter values can be uniformly updated during each round of training, thereby ensuring that the training accuracy of the AI model can reach a high level.
[0057] Furthermore, during the iterative training of the AI model, the computing node 101 will store the parameter values obtained through different rounds of training in different storage areas in the memory 1012. There is no need to perform the operation of copying the parameter values in the memory 1012, which can avoid the impact of the copied parameter values on the efficiency of the AI model training of the computing node 101.
[0058] In actual applications, during the training of a large AI model with a large number of parameters (such as trillions of parameters), the utilization rate of the memory 1012 in the computing node 101 is usually small, such as less than 50%. Therefore, the existing storage space in the memory 1012 can also support the computing node 101 to store the parameter values obtained from 2 (or more than 2) rounds of training, and no additional hardware requirements are required for the distributed training system 20.
[0059] It can be understood that the above is an introduction to the process of training the AI model based on computing node 101. For other computing nodes in the distributed training system 20, similar methods can also be used to train and (after failure) restore the AI model, which will not be elaborated on.
[0060] As an example, the distributed training system 20 shown in FIG. 2 can be applied to the training scenario shown in FIG. 3 .
[0061] As shown in Figure 3, the computing nodes in the distributed training system 20 can specifically be accelerator cards in the training server, such as computing node 101 can be accelerator card 3111 in training server 31, etc.; the switching nodes can specifically be switches, such as switching node 201 can be switch 401, etc.
[0062] As shown in Figure 3 , each training server may include not only an accelerator card but also a CPU. For example, training server 31 includes CPU 301. This CPU can manage the accelerator card, such as controlling the accelerator card to start or stop training an AI model. The CPU in the training server can be configured with memory and a network card. As shown in Figure 3 , CPU 301 can be configured with a network card 3012 and memory 3013.
[0063] Among them, the CPU in the training server can receive the data set sent from the data plane network shown in Figure 3 through the network card, and provide the data set to the accelerator card. The data set is used to train the AI model on the accelerator card. Different accelerator cards can obtain the same or different data sets. In addition, the CPU can also receive control commands sent from the management plane network shown in Figure 3 through the network card, such as receiving control commands sent from the scheduling server 33 in the management plane network, so as to instruct the accelerator card to start the training process for the AI model according to the control command. Thus, each accelerator card can use the data set forwarded by the CPU to iteratively train the AI model deployed on the accelerator card.
[0064] Furthermore, if multiple accelerator cards train the same AI model in parallel, during each round of iterative training, each accelerator card can obtain the parameter values trained by other accelerator cards through the network card and the parameter plane network shown in Figure 3. After aggregating the parameter values trained by multiple accelerator cards, the aggregated results are sent to other accelerator cards through the parameter plane network, and gradient synchronization is performed with other accelerator cards, thereby enabling multiple accelerator cards to synchronously update the parameter values of the same AI model. In actual applications, accelerator cards located in the same training server can communicate through the parameter plane network or through an internal bus, which is not limited to this.
[0065] Alternatively, when multiple accelerator cards train the same AI model in parallel, during each round of iterative training, each accelerator card can send the parameter values obtained through training to the parameter plane network, so that the switches in the parameter plane network (such as switch 403, etc.) can aggregate the parameter values obtained through training of multiple accelerator cards and send the aggregation results to each accelerator card (i.e., in-network computing), thereby achieving synchronous update of the parameter values in the AI models on multiple accelerator cards.
[0066] In actual application, the management plane network, data plane network and parameter plane network shown in Figure 3 can share the same network, or they can be implemented through independent networks respectively. Figure 3 uses the example of the management plane network and the data plane network sharing the same network and the parameter plane network using a separate network, and this is not limited to this.
[0067] It is worth noting that in addition to the training scenario shown in Figure 3 above, the distributed training system 20 shown in Figure 2 can also be applied to other training scenarios. For example, in other possible training scenarios, the computing nodes in the distributed training system 20 shown in Figure 2 can be CPUs in the training server, so that the AI model can be distributedly trained by multiple CPUs.
[0068] In addition, the above-mentioned figure 2 and figure 3 introduce the process of distributed training of the AI model. In other application scenarios, a single computing node can also be used to train an AI model (the number of parameters of the AI model can be in the order of hundreds of thousands, etc.). Moreover, during the iterative training of the AI model, the computing node can alternately save the parameter values generated during multiple rounds of training in different storage areas in the memory, and persist the parameter values in the memory after a failure occurs to facilitate the subsequent recovery of the AI model.
[0069] For ease of understanding, the following describes an embodiment of the AI model training method provided in this application in conjunction with the accompanying drawings.
[0070] Referring to Figure 4, Figure 4 is a flow chart illustrating an AI model training method provided in an embodiment of the present application. This method can be applied to the distributed training system shown in Figure 2 or Figure 3, or to other applicable distributed training systems. For ease of illustration, this embodiment uses the distributed training system 20 shown in Figure 2 as an example to describe the process of iteratively training an AI model by computing node 101.
[0071] As shown in FIG4 , the model training method may specifically include:
[0072] S401: During the Nth round of training of the AI model, the computing node 101 stores the state data obtained through training to the storage area 1 in the memory 1012. The state data includes at least the parameter values of the AI model after the Nth round of training, where N is a positive integer.
[0073] In this embodiment, multiple computing nodes in the distributed training system 20 can jointly train the same AI large model. At this time, the AI model trained by the computing node 101 can be a complete AI large model (such as the AI large model includes a small number of parameters). The distributed training system 20 can train the AI large model in a data parallel manner, that is, the same AI model can be deployed on the computing node 101 and the computing node 104, and the AI model on each computing node is a complete AI large model, so that different computing nodes can use different data sets to train the AI model in parallel. Accordingly, in each round of training the AI model, each computing node can use the aggregation result to update the parameter value. The aggregation result can be obtained by aggregating the parameter values trained by each computing node, which can be aggregated by the computing node or by the switch in the parameter plane network (i.e., in-network computing). Exemplarily, the parameter plane network can be a remote direct memory access over converged ethernet (RoCE) network on a converged Ethernet network, or an IB network (InfiniBand network), etc.
[0074] Alternatively, the AI model trained by computing node 101 may be a partial model within a large AI model (e.g., an AI model that includes a large number of parameters), such as a partial network layer within the large AI model. In this case, the distributed training system 20 may train the AI model in a model-parallel manner, such as using computing node 101 to train a partial model within the large AI model, computing node 104 to train another partial model within the large AI model, and different computing nodes to train different parts of the AI model.
[0075] As an implementation example, during the Nth round of training of the AI model, the computing node 101 can input the input data in the training sample 1 into the AI model, and the AI model performs reasoning based on the input data to obtain the corresponding reasoning result. Then, the computing node 101 can compare the difference between the reasoning result and the label in the training sample 1 (usually the real result). Thus, the computing node 101 can calculate the gradient used to update the parameter value in the AI model based on the difference between the reasoning result and the label, and further update the parameter value accordingly based on the gradient, thereby completing a round of training process for the AI model. Among them, when the AI model trained by the computing node 101 is a complete model, each time the computing node 101 updates the parameter value, it can first aggregate the gradients calculated by multiple computing nodes to obtain a gradient aggregation result, and then use the gradient aggregation result to update the parameter value.
[0076] Typically, during the AI model training process, the computing node 101 stores the generated data in the memory 1012. The generated data may include activation values, parameter values, gradients, momentum, and so on. Activation values refer to the output values generated by the network layer in the AI model based on the inputs. Parameter values refer to the values of model parameters, such as weights, during the AI model training process. The values of model parameters are updated during each round of training. Each parameter value can be 16-bit binary floating point precision (FP16) or 32-bit binary floating point precision (FP32). Gradients refer to the gradients used to update the values of model parameters and can indicate the magnitude of the change in the values of the model parameters in a certain direction. Momentum refers to the value used to adjust the gradient change. The gradient value used to update the parameter value can be adjusted based on the gradient value used when the parameter value was previously updated, thereby improving the stability of model training. For example, momentum in the AI model can include first-order momentum, second-order momentum, and so on, but this is not limited to this.
[0077] In one possible implementation, the memory 1012 may include three storage areas, namely storage area 1, storage area 2, and storage area 3. During each round of AI model training, the computing node 101 may store state data in the generated data in storage area 1 and store other data in storage area 2. State data refers to data that can indicate the state of the AI model, or data that can be used to restore the AI model.
[0078] Exemplarily, the state data may be parameter values, specifically parameter values in the AI model after completing the Nth round of training. Accordingly, other data stored in storage area 2 may include momentum (such as first-order momentum and second-order momentum, etc.), activation values and gradients generated in the Nth round of training, and other data. In other examples, the state data may include not only parameter values, but also other data that can be used to restore the state of the AI model. For example, the state data may include parameter values, momentum, etc. At this time, the data stored in storage area 2 include activation values and gradients.
[0079] In actual application, when the storage space in the memory 1012 is large enough (such as greater than a threshold), the computing node 101 uses the state data stored in the storage area 1 to include not only parameter values but also other data (such as momentum); and when the storage space in the memory 1012 is small (such as less than the threshold), the computing node 101 uses the state data stored in the storage area 1 to include only parameter values. Alternatively, when the storage space in the memory 1012 is large enough, during the model training process, the computing node 101 can generate high-precision parameter values, such as parameter values with 32-bit floating point precision, and accordingly, the computing node 101 can store the high-precision parameter values in the storage area 1; and when the storage space in the memory 1012 is small, the computing node 101 can generate lower-precision parameter values based on the high-precision parameter values, such as parameter values with 16-bit floating point precision, and store the lower-precision parameter values in the storage area 1, thereby reducing the memory resources required to store the parameter values.
[0080] S402: During the N+1th round of training of the AI model, the computing node 101 gradually updates the parameter values in the AI model.
[0081] Typically, after completing the Nth round of model training, the computing node 101 may determine whether the AI model meets the model training termination conditions, such as whether the number of iterative training of the AI model has reached a preset number, or whether the AI model has converged. If the termination conditions of the model training are not met, the computing node 101 may continue to perform the next round of training on the AI model, i.e., the N+1th round of training.
[0082] Similar to the Nth round of AI model training, during the N+1th round of training, computing node 101 can input the input data in training sample 2 into the AI model, and the AI model will perform inference based on the input data to obtain a corresponding inference result, and compare the difference between the inference result and the label in training sample 2. Thus, computing node 101 can calculate the gradient used to update the parameter value in the AI model based on the difference between the inference result and the label, and further update the parameter value accordingly based on the gradient. In this embodiment, it is assumed that the update of some parameter values has been completed, and the model parameters that have been updated are the first model parameters described in step S402.
[0083] Among them, for multiple parameters in the AI model, the computing node 101 can gradually update the values of the multiple parameters. For example, when the AI model is specifically a neural network model, the AI model may include K network layers, where K is a positive integer greater than 1. Among them, each network layer may include multiple parameters. Then, when updating the parameter values in the AI model, the computing node 101 may first update the values of the parameters in the K-th layer network. After the values of all parameters in the K-th layer network are updated, the values of the parameters in the (K-1)-th layer network are updated, and then the values of the parameters in the (K-2) layer are updated, and so on, until the parameter values in the 1st layer network are updated.
[0084] S403 : The computing node 101 stores the updated parameter value in the storage area 3 in the memory 1012 .
[0085] In this embodiment, when updating the parameter values in the AI model, the computing node 101 does not replace the parameter values before the update stored in the storage area 1 with the updated parameter values. Instead, the updated parameter values are stored in a new storage area (i.e., storage area 3) included in the memory 1012, thereby preventing the parameter values obtained after the Nth round of model training from being overwritten. In this way, even if a failure occurs during the process of updating the model parameter values in the N+1th round, the parameter values stored in the storage area 1 in the memory 1012 are the values of all the parameters of the AI model after the Nth round of training, and there will be no situation where some parameter values are the values after the Nth round of model training, while other parameter values are the updated values generated by the N+1th round of model training.
[0086] As for activation values, gradients and other data, during the N+1 round of training, computing node 101 can save the newly generated activation values, gradients and other data through storage area 2, specifically replacing the old activation values and gradient data in storage area 2 with the newly generated activation values and gradient data.
[0087] S404: When a failure occurs during the N+1 round of training of the AI model, the computing node 101 starts a fault saving mechanism, which includes the status data in the storage area 1 in the persistent storage memory 1012.
[0088] In actual applications, during the training of AI models, failures are inevitable, such as interruptions caused by program errors in computing node 101, or interruptions caused by abnormal communication functions in computing node 101 (in data parallel scenarios). If computing node 101 fails while updating the parameter values of the AI model after completing the update of some parameter values and continuing to update the remaining parameter values, the model training of the N+1 round will fail. At this time, since storage area 1 stores the complete model parameter values generated by the AI model after the Nth round of training (which may also include other data such as momentum), computing node 101 can activate a fault preservation mechanism to persistently store the state data in storage area 1 in memory 1012, so that the currently trained AI model can be restored to the AI model obtained after the Nth round of training using this persistently stored state data. In actual applications, after activating the fault preservation mechanism, computing node 101 can also persistently store other data, such as activation values, gradients, and other data in storage area 2, or can persistently store files of the AI model, etc., without limitation.
[0089] In this way, after completing the fault recovery, the computing node 101 can re-perform the N+1 round of training on the AI model based on the persistently stored state data. In this way, in which round of model training does a fault occur, the computing node 101 can continue to train the AI model from this round, which enables the training loss time caused by the fault to be controlled within a range less than the time required for one round of AI model training, as shown in Figure 5. In actual application, the training loss time can be reduced to tens of seconds or even a few seconds, thereby effectively reducing the training loss time to improve the overall efficiency of training the AI model and improve the availability of the distributed training system 20.
[0090] For example, when a failure occurs, computing node 101 can persistently store the state data in storage area 1 to the local hard disk of computing node 101, or persistently store the state data in cloud storage mounted to computing node 101. Accordingly, after the failure is recovered, computing node 101 can obtain the state data from the local hard disk or cloud storage, and restore the AI model to the AI model obtained after completing the Nth round of training based on the state data. Thus, computing node 101 can continue to perform the N+1th round of training on the AI model.
[0091] Furthermore, when the AI model trained by computing node 101 is a complete model in the AI large model required to be trained by the distributed training system 20, since multiple computing nodes (including computing node 101) will train the same AI model in parallel, and the same state data will be saved in memory during each round of training of the AI model. Therefore, after the failure is recovered, computing node 101 can not only restore the AI model from the state data persistently stored in the local hard disk or cloud storage, but also obtain the state data by accessing the memory of other computing nodes, and restore the AI model on computing node 101 to the AI model obtained after completing the Nth round of training based on the state data in the memory of other computing nodes, so that computing node 101 can continue to perform the N+1th round of training on the AI model. Moreover, since multiple computing nodes train the AI model synchronously, other computing nodes except computing node 101 can also restore the AI model to the AI model obtained after completing the Nth round of training based on the state data saved in their respective memories, so that multiple computing nodes can continue to perform the N+1th round of training on the AI model.
[0092] Among them, the multiple computing nodes in the distributed training system 20 can be under the control of the scheduling server (such as the scheduling server 33 shown in Figure 3) to perform the process of training the AI model. Specifically, before the computing node 101 fails, the scheduling server can send control commands to each computing node respectively to control each computing node to execute the model training process. In addition, the scheduling server can sense that the AI model has failed, such as when the computing node 101 sends a fault indication information to the scheduler when the failure occurs, or the scheduling server does not receive the heartbeat message sent by the computing node 101 for a long time. Then, when the computing node 101 completes the fault recovery, the scheduling server can send a new control command to the computing node 101 (and other computing nodes) to instruct the computing node 101 (and other computing nodes) to resume training for the AI model.
[0093] In this embodiment, two rounds of training process for the AI model are used as an example for explanation. During the training process before the Nth round of the AI model and the training process after the N+1th round, the state data generated by the previous round of training and the model state number generated by the current round of training can be saved in different storage areas in the memory 1012 in a similar manner as described above, so that when the computing node 101 fails, the AI model can be restored to the AI model after the completion of the previous round of training.
[0094] For example, since the storage space of memory 1012 is usually limited, it is difficult to support the computing node 101 to save the state data generated during each round of AI model training in a separate storage area. Therefore, the computing node 101 can alternately store the state data generated during two adjacent rounds of model training through storage area 1 and storage area 3. For example, during the Nth round of training, the computing node 101 can save the state data through storage area 1; during the N+1th round of training, the computing node 101 can save the new state data through storage area 3; during the N+2th round of training, the computing node 101 uses storage area 1 to save the new state data, that is, the state data generated during the N+2th round of training is used to overwrite the state data generated during the Nth round of training.
[0095] Furthermore, the timing of a failure of computing node 101 is not limited to the process of updating the second parameter value. For example, if computing node 101 fails before updating the first parameter value, computing node 101 can also resume training of the AI model using the state data stored in storage area 1.
[0096] In the embodiment shown in FIG4 , the AI model trained by computing node 101 is a portion of a complete large AI model, such as a portion of a network layer. Alternatively, the AI model trained by computing node 101 can be a complete large AI model, i.e., different computing nodes train the same large AI model using different training samples. These two training scenarios are exemplified below with reference to the accompanying figures.
[0097] Referring to Figure 6 , a flow chart of a method for training an AI model is shown. In the embodiment shown in Figure 6 , the AI large model to be trained in the distributed training system 20 includes four AI models, namely AI model 1, AI model 2, AI model 3, and AI model 4. Furthermore, the four AI models are trained by computing nodes 101 to 104, respectively. Each computing node is used to train one AI model, and different AI models can be different parts of the AI large model.
[0098] Taking computing node 101 training AI model 1 as an example, as shown in FIG6 , the method may specifically include:
[0099] S601: During the Nth round of training of the AI model 1, the computing node 101 stores the state data obtained through training to the storage area 1 in the memory 1012. The state data includes parameter values, and the activation values, gradients, and momentum generated during the model training process are stored to the storage area 2 in the memory 1012.
[0100] S602: During the N+1 round of training of AI model 1, computing node 101 stores the newly generated activation value, gradient, and momentum into storage area 2.
[0101] S603: The computing node 101 gradually updates the parameter values in the AI model 1.
[0102] S604: The computing node 101 stores the updated parameter value in the storage area 3.
[0103] The specific implementation process of steps S601 to S604 can be found in the relevant description of steps S401 to S403 in the embodiment shown in FIG4 , and will not be repeated here. Furthermore, during the N+1 round of training of AI model 1, computing node 101 can concurrently perform operations of storing gradients in storage area 2 and storing updated parameter values in storage area 3.
[0104] S605: When a failure occurs during the N+1th round of training of AI model 1, computing node 101 persistently stores the status data in storage area 1.
[0105] In this embodiment, when a fault occurs during the process of gradually updating the parameter values in the AI model, the computing node 101 can start a fault saving mechanism, which may include persistent storage of status data in the storage area 1, and may also include continuing to iteratively train the AI model that has undergone the Nth round of training after the computing node 101 completes fault recovery.
[0106] For example, the failure of the computing node 101 may be, for example, a failure of the program code for training the AI model running on the computing node 101.
[0107] As shown in FIG7 , the computing node 101 can back up the state data in the storage area 1 to a persistent storage medium, and the state data can be used as a checkpoint for restoring model training.
[0108] S606: After the computing node 101 completes fault recovery, the computing node 101 can restore the AI model 1 to the AI model 1 after completing the Nth round of training based on the persistently stored state data.
[0109] Illustratively, computing node 101 can complete fault recovery by executing a preset fault handling strategy. For example, computing node 101 can complete fault recovery by restarting computing node 101. Alternatively, computing node 101 can complete fault recovery under the operation and maintenance of an operation and maintenance personnel. For example, the operation and maintenance personnel can upgrade the software of computing node 101, so that computing node 101 can resume normal model training function by running the upgraded software.
[0110] As shown in FIG7 , computing node 101 can read the state data stored in the persistent storage medium into memory 1012 , for example, by reading the state data into storage area 1. Computing node 101 can then use the state data in storage area 1 to restore AI model 1 to the AI model 1 obtained when the Nth round of training is completed. Furthermore, when continuing to perform the N+1th round of training on AI model 1, when updating parameters in AI model 1, the momentum in the state data can be used to calculate the gradient for updating the parameter, so that the parameter value is updated according to the calculated gradient.
[0111] Furthermore, when continuing to perform the N+1th round of training on AI model 1, the newly generated state data (including the newly generated parameter values and the new momentum) can be saved to the storage area 3 in the memory 1012.
[0112] S607: Computing node 101 continues to train AI model 1 until the training termination condition is met.
[0113] Exemplarily, the training termination condition may be, for example, that the number of iterative training times of the AI model 1 meets a preset number, or that the AI model 1 is in a convergence state, etc., which is not limited.
[0114] Similarly, for the process of training AI models for computing nodes 102 to 104 respectively, please refer to the above description of the relevant parts about training AI model 1 for computing node 101, which will not be repeated here.
[0115] Referring to Figure 8, a flow chart of a method for AI model training is shown. In the embodiment shown in Figure 8, computing nodes 101 to 104 use different training samples to train the same AI model (i.e., the large AI model required to be trained by the distributed training system 20), and in each round of model training, computing node 101 aggregates the parameter values obtained by training multiple computing nodes and feeds the aggregated results back to the remaining computing nodes, so that each computing node uses the aggregated set to update the parameters of the AI model.
[0116] Taking computing node 101 training an AI model as an example, as shown in FIG8 , the method may specifically include:
[0117] S801: During the Mth round of training of the AI model, the computing node 101 stores the state data obtained through training to the storage area 1 in the memory 1012. The state data includes the parameter values in the AI model, and stores the activation value, gradient, and momentum generated during the model training process to the storage area 2 in the memory 1012. M is a positive integer.
[0118] S802: During the M+1 round of training of the AI model, the computing node 101 stores the newly generated activation value, gradient, and momentum into storage area 2.
[0119] S803: The computing node 101 gradually updates the parameter values in the AI model.
[0120] S804: The computing node 101 stores the updated parameter value in the storage area 3.
[0121] Since multiple computing nodes 101 train the same AI model in parallel, each time computing node 101 updates the first parameter value, it first synchronizes the gradient with the other computing nodes. This means that the gradient used to update the first parameter value is synchronized with the remaining computing nodes. After gradient synchronization is completed, each computing node (including computing node 101) uses the gradient to update the first parameter value, as shown in FIG9 .
[0122] S805: When a failure occurs during the M+1th round of training of the AI model, the computing node 101 stops training the AI model.
[0123] Normally, since the pace at which the remaining computing nodes train the AI model is consistent with the pace at which computing node 101 trains the AI model, when computing node 101 fails and stops training the AI model, the remaining computing nodes can also suspend the training process for the AI model, thereby suspending the model training process of the distributed training system 20, as shown in Figure 9.
[0124] S806: The computing node 101 determines whether a hardware failure occurs. If so, the process proceeds to step S807; if not, the process proceeds to step S809.
[0125] S807: After completing fault recovery, computing node 101 obtains the status data of the AI model from the memory of computing node 102. The status data is consistent with the status data written to storage area 1 by computing node 101 when completing the Mth round of training.
[0126] A hardware failure, for example, could be a hardware failure on compute node 101, such as physical damage to the processor, memory, or network card in compute node 101. In this case, operations and maintenance personnel can recover from the failure by restarting or replacing the faulty hardware on compute node 101. Because multiple compute nodes train AI models in unison, the state data stored in memory across these nodes is typically consistent. This means that even if a failure in compute node 101 causes the loss of state data in storage area 1, compute node 101 can still retrieve the same state data from the memory of other compute nodes.
[0127] S808: The computing node 101 restores the AI model to the AI model after completing the Mth round of training based on the state data. Then, the computing node 101 proceeds to step S811.
[0128] Among them, since the parameter values are updated based on gradient synchronization between different computing nodes, when computing node 101 restores the AI model to the AI model after completing the Mth round of training, the remaining computing nodes 101 can also restore their own trained AI models to the AI model after completing the Mth round of training based on the status data stored in the memory, as shown in Figure 9, so that each computing node can continue to train the AI model in parallel based on the same pace.
[0129] S809: The computing node 101 persistently stores the state data stored in the storage area 1.
[0130] S810: After the computing node 101 completes fault recovery, the computing node 101 can restore the AI model to the AI model after completing the Mth round of training based on the persistently stored state data.
[0131] S811: Computing node 101 continues to train the AI model until the training termination condition is met.
[0132] In this embodiment, the example of continuing to train the AI model on the computing node 101 after the computing node 101 has recovered from the failure is used for illustration. In other embodiments, when a hardware failure occurs in the computing node 101, the computing node can also be replaced. For example, the operation and maintenance personnel can deploy a new computing node in the distributed training system 20 and resume the training of the AI model on the replaced computing node.
[0133] It is worth noting that other reasonable step combinations that can be thought of by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be familiar with that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0134] The above introduces the distributed training system and AI model training method provided in the embodiment of the present application in combination with Figures 1 to 9. Next, the structure of the computing device for implementing the computing node 101 provided in the embodiment of the present application is introduced in combination with the accompanying drawings.
[0135] 10 , which shows a schematic diagram of the hardware structure of a computing device 1000 . The computing device 1000 , for example, can implement the computing node 101 in the embodiment shown in FIG. 2 .
[0136] As shown in Figure 10, the computing device 1000 includes a processor 1001, a memory 1002, and a communication interface 1003. The processor 1001, the memory 1002, and the communication interface 1003 communicate through a bus 1004, and may also communicate through other means such as wireless transmission. The memory 1002 is used to store instructions, and the processor 1001 is used to execute the instructions stored in the memory 1002. Furthermore, the computing device 1000 may also include a memory unit 1005, which may be used to cache program code, or may be used to cache data, such as the above-mentioned status data. The memory unit 1005 may be connected to the processor 1001, the storage medium 1002, and the communication interface 1003 via a bus 1004. The memory 1002 stores program code, and the processor 1001 may call the program code stored in the memory 1002 to perform the following operations:
[0137] During the Nth round of training of the AI model, the first computing node stores first state data of the AI model in a first storage area in the memory, where the first state data includes parameter values of the AI model after the Nth round of training, and the AI model is hosted on the first computing node, where N is a positive integer;
[0138] During the N+1th round of training of the AI model, the first computing node gradually updates the parameter values in the AI model and stores the updated parameter values in the second storage area in the memory;
[0139] When a failure occurs during the N+1th round of training of the AI model, the first computing node starts a fault saving mechanism, which includes persistently storing the first state data in the first storage area in the memory.
[0140] It should be understood that in this embodiment, the processor 1001 may be a CPU, or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete device components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0141] The memory 1002 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1001. The memory 1002 may also include a nonvolatile random access memory.
[0142] The memory 1002 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0143] The communication interface 1003 is used to communicate with other devices connected to the computing device 1000. The bus 1004 may include, in addition to the data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all buses are labeled as bus 1004 in the figure.
[0144] It should be understood that the computing device 1000 according to the embodiment of the present application may correspond to the method executed by the computing node 101 in the method shown in Figure 4, Figure 6 or Figure 8 in the embodiment of the present application. The above-mentioned and other operations and / or functions implemented by the computing device 1000 are respectively for implementing the process of the corresponding method in Figure 4, Figure 6 or Figure 8. For the sake of brevity, they will not be repeated here.
[0145] The present application provides an accelerator card, comprising a power supply circuit and a processing circuit. The power supply circuit is used to supply power to the processing circuit, and the processing circuit is used to execute the method executed by the computing node 101 in the method shown in FIG4, FIG6 or FIG8.
[0146] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned AI model training method.
[0147] The present application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the computer program product fully or partially generates the process or function described in the present application.
[0148] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0149] The computer program product may be a software installation package, which may be downloaded and executed on a computing device when any of the aforementioned model training methods is required.
[0150] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0151] The terms used in the above embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of this application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; the character " / " generally indicates that the objects associated with each other are in an "or" relationship. In the embodiments of the present application. "Simultaneously" means within the same time period, including situations at the same time. The terms "first", "second", etc. in the specification, claims and drawings of this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate, and this is merely a way of distinguishing objects with the same properties when describing them in the embodiments of the present application.
[0152] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0153] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. An artificial intelligence AI model training method, characterized in that, The method includes: During the Nth round of training the AI model, the first computing node stores the first state data of the AI model in the first storage area in the memory. The first state data includes the parameter values of the AI model after the Nth round of training. The AI model is carried on the first computing node, and N is a positive integer. During the (N + 1)th round of training the AI model, the first computing node gradually updates the parameter values in the AI model and stores the updated parameter values in the second storage area in the memory. When a failure occurs during the (N + 1)th round of training the AI model, the first computing node activates a fault preservation mechanism. The fault preservation mechanism includes persistently storing the first state data in the first storage area in the memory.
2. The method according to claim 1, wherein The fault preservation mechanism further includes: After the first computing node completes fault recovery, based on the persistently stored first state data, the AI model after the Nth round of training is restored. Continue to perform the (N + 1)th round of training on the AI model.
3. The method according to claim 1 or 2, characterized in that, The first computing node belongs to a distributed training system. The distributed training system includes multiple computing nodes. The multiple computing nodes are used for jointly training an AI large model. The multiple computing nodes include the first computing node. Among them, the AI model is part or all of the AI large model.
4. The method according to claim 3, wherein The multiple computing nodes further include a second computing node. The second computing node is configured with the same AI model as the first computing node. The method further includes: When a failure occurs during the (M + 1)th round of training the AI model, the first computing node obtains second state data from the second computing node. The second state data includes the parameter values of the AI model after the Mth round of training. M is a positive integer. After the first computing node completes fault recovery, the first computing node restores the AI model after the Mth round of training according to the second state data. The first computing node continues to perform the (M + 1)th round of training on the AI model.
5. The method according to claim 3, characterized in that, The multiple computing nodes further include a third computing node. The AI model carried on the first computing node is part of the AI large model, and the AI model carried on the third computing node is another part of the AI large model.
6. The method according to any one of claims 3 to 5, characterized in that, The failure is a failure that occurs in the software program participating in training the AI model on the first computing node, or the failure is a communication failure that occurs in the first computing node.
7. The method according to any one of claims 1 to 6, characterized in that, The first state data further includes momentum, and the momentum is used to update the parameter values of the AI model during the (N + 1)th round of training the AI model.
8. The method according to any one of claims 1 to 6, characterized in that, The first computing node is an acceleration card, and the memory is a high-bandwidth memory HBM.
9. A distributed training system, characterized in that, The distributed training system includes multiple computing nodes. The multiple computing nodes are used for jointly training an AI large model. The multiple computing nodes include a first computing node. The AI model carried on the first computing node is part or all of the AI large model. The first computing node is configured to: During the Nth round of training of the AI model, store the first state data of the AI model in a first storage area in the memory. The first state data includes the parameter values of the AI model after the Nth round of training. The AI model is carried on the first computing node, and N is a positive integer. During the (N + 1)th round of training of the AI model, gradually update the parameter values in the AI model and store the updated parameter values in a second storage area in the memory. When a fault occurs during the (N + 1)th round of training of the AI model, start a fault saving mechanism, which includes persistently storing the first state data in the first storage area in the memory.
10. The distributed training system according to claim 9, wherein The fault saving mechanism further includes: After the first computing node completes fault recovery, based on the persistently stored first state data, recover the AI model after the Nth round of training is completed. Continue to perform the (N + 1)th round of training on the AI model.
11. The distributed training system according to claim 9, wherein The multiple computing nodes further include a second computing node, and the second computing node is configured with the same AI model as the first computing node. The first computing node is further configured to: When a fault occurs during the (M + 1)th round of training of the AI model, obtain second state data from the second computing node. The second state data includes the parameter values of the AI model after the Mth round of training, and M is a positive integer. After the first computing node completes fault recovery, based on the second state data, recover the AI model after the Mth round of training is completed. Continue to perform the (M + 1)th round of training on the AI model.
12. The distributed training system according to claim 9, wherein The multiple computing nodes further include a third computing node. The AI model carried on the first computing node is a partial model of the AI large model, and the AI model carried on the third computing node is another partial model of the AI large model.
13. The distributed training system according to any one of claims 9 to 12, characterized in that The fault is a fault that occurs in the software program participating in the training of the AI model on the first computing node, or the fault is a communication fault that occurs in the first computing node.
14. The distributed training system according to any one of claims 9 to 13, characterized in that The first state data further includes momentum, and the momentum is used to update the parameter values of the AI model during the (N + 1)th round of training of the AI model.
15. The distributed training system according to any one of claims 9 to 14, characterized in that, The first computing node is an acceleration card, and the memory is a high-bandwidth memory HBM.
16. An acceleration card, characterized in that, The acceleration card is used to execute the method according to any one of claims 1 to 8.
17. A computing device, characterized in that, The computing device includes a processor and a memory. The memory is used to store instructions, and the processor executes the instructions stored in the memory so that the computing device executes the method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that, Includes instructions that, when running on a computing device, cause the computing device to execute the method according to any one of claims 1 to 8.
19. A computer program product comprising instructions, characterized in that, When running on a computing device, cause the computing device to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
AI model training method, distributed training system and related equipment
CN120373489A
Distributed training method, device and system of model
CN113094168A
Model training method and device, equipment and storage medium
CN114298329A
Method and device for training machine learning model executed by using parameter server
CN116414615A
Model training method and device, storage medium and electronic equipment
CN116755941A
Cited By
Distributed large model
CN121212401A
Data processing method and device, terminal equipment and storage medium
CN121303365A