Data processing method and computer equipment
By deploying multiple copies of optimizer data in the computing system and transmitting gradients between processors, the problem of processor failures resulting in incomplete optimizer data is solved, and the integrity of optimizer data and the efficiency of model training is achieved.
Patent Information
- Application Number
- CN202311690461.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
When multi-processors collaboratively train artificial intelligence models, processor failures may lead to incomplete optimizer data, resulting in inaccurate checkpoint information, and wasting model training time.
By deploying at least two copies of optimizer data in a computing system and transmitting gradients between dedicated processors, the optimizer data is updated to ensure that the complete optimizer data can still be formed in the event of a processor failure.
The completeness and reliability of the optimizer data is achieved, the accuracy of end-of-life checkpoints is ensured, the loss of model training is reduced and the training efficiency is improved.
Smart Images

Figure CN120123142A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a data processing method and a computer device. Background Art
[0002] Currently, multiple processors cooperate to train an artificial intelligence (AI) model with a large number of parameters. For example, a large language model (LLM) is trained based on multi-machine multi-card distributed training, and the training duration of the LLM can reach several months. To prevent the interruption of model training caused by failures of processors, networks, software, etc., the information of checkpoints during the model training process can be saved periodically. For example, the information of checkpoints includes model parameters and optimizer data.
[0003] However, each of the multiple processors saves a part of the optimizer data, and the parts saved by each processor are different. There is only one copy globally, and the data saved by the multiple processors can form complete optimizer data. If a processor fails, some of the optimizer data stored by the processor may be lost, resulting in incomplete optimizer data, inaccurate checkpoint information, and wasting the time of model training. Summary of the Invention
[0004] This application provides a data processing method and a computer device, thereby ensuring the integrity of optimizer data when a processor fails.
[0005] In a first aspect, a data processing method is provided. The computing system to which the method is applied includes a general-purpose processor and multiple dedicated processors, and the multiple dedicated processors are used to train an artificial intelligence model. Among them, the multiple dedicated processors include at least two copies of optimizer data, and the optimizer data includes a first data set, and at least two dedicated processors both include the first data set. For example, the multiple dedicated processors include a first dedicated processor and a second dedicated processor, and both the first dedicated processor and the second dedicated processor include the first data set. The method includes: the first dedicated processor trains the artificial intelligence model to obtain a first gradient; the first dedicated processor updates the first data set according to the first gradient to obtain an updated first data set; the second dedicated processor updates the first data set according to the first gradient to obtain an updated first data set. When the second dedicated processor fails, the general-purpose processor instructs to persist the data sets included in the multiple non-failed dedicated processors, and the data sets included in the multiple non-failed dedicated processors form the updated optimizer data. The multiple non-failed dedicated processors include the first dedicated processor, and the updated optimizer data includes the updated first data set.
[0006] Compared with a single copy of optimizer data globally, when a dedicated processor fails, it may lead to inaccurate end-of-life checkpoints and incomplete optimizer data. In the method provided by this application, at least two dedicated processors globally contain data sets in the optimizer data, that is, at least two copies of optimizer data are deployed globally to achieve multi-copy optimizer data globally. Moreover, the dedicated processors update the data sets in the optimizer data they contain by transmitting gradients, that is, update the multi-copy optimizer data by "computing instead of transmitting", ensuring the integrity of the optimizer data at any time. When a dedicated processor fails, since the system contains multiple copies of optimizer data, a complete optimizer data can be composed of the data sets contained in the non-failed dedicated processors, improving the reliability of the optimizer data, ensuring accurate end-of-life checkpoints, facilitating the restoration of model training based on the end-of-life checkpoints, and reducing training losses.
[0007] In a possible implementation, the method further includes: the first dedicated processor obtains the initial value of the first data set configured by the general processor; the second dedicated processor obtains the initial value of the first data set sent by the first dedicated processor.
[0008] Thus, it is ensured that at least two dedicated processors both contain the same data set.
[0009] In another possible implementation, the method further includes: the first dedicated processor transmits the first gradient to the second dedicated processor.
[0010] Thus, it is convenient for the dedicated processors containing the same data set to obtain the gradient and update the data set, ensuring that at least two dedicated processors also contain the same updated data set. Since the amount of optimizer data is large and the amount of gradient data is small, transmitting gradients can reduce the bandwidth and the amount of data transmitted.
[0011] In another possible implementation, the optimizer data further includes a second data set, and multiple dedicated processors include a third dedicated processor and a fourth dedicated processor, and both the third dedicated processor and the fourth dedicated processor contain the second data set.
[0012] Understandably, the first data set contained in the first dedicated processor and the second data set contained in the third dedicated processor constitute complete optimizer data. The first data set contained in the second dedicated processor and the second data set contained in the fourth dedicated processor constitute complete optimizer data.
[0013] Thus, at least two dedicated processors globally contain data sets in the optimizer data, that is, at least two copies of optimizer data are deployed globally to achieve multi-copy optimizer data globally.
[0014] In another possible implementation, the method further includes: a third dedicated processor training an artificial intelligence model to obtain a second gradient; the third dedicated processor updating a second data set according to the second gradient to obtain an updated second data set; a fourth dedicated processor obtaining the second gradient from the third dedicated processor; the fourth dedicated processor updating the second data set according to the second gradient to obtain an updated second data set.
[0015] Thus, by "replacing transmission with computing", the optimizer data of multiple replicas is updated, ensuring the integrity of the optimizer data at any time.
[0016] In another possible implementation, the general-purpose processor instructs to persist the data sets included in multiple non-faulty dedicated processors, including: obtaining integrity information sent by non-faulty dedicated processors in the computing system, where the integrity information is used to indicate the integrity of the data sets included in the non-faulty dedicated processors, determining, according to the integrity information sent by non-faulty dedicated processors in the computing system, that the data sets included in multiple non-faulty dedicated processors form updated optimizer data, and sending a persistence command to multiple non-faulty dedicated processors. The persistence command is used to instruct to persist the data sets included in multiple non-faulty dedicated processors.
[0017] In this way, the general-purpose processor determines the dedicated processors that can form complete optimizer data based on the integrity information of the data included in the dedicated processors, ensuring the integrity of the optimizer data when a dedicated processor fails, improving the reliability of the optimizer data, ensuring the accuracy of the end-of-life checkpoint, facilitating the restoration of model training based on the end-of-life checkpoint, and reducing the training loss.
[0018] In another possible implementation, the integrity information includes the identifiers and correct information of the data in the data set; determining, according to the integrity information sent by non-faulty dedicated processors in the computing system, that the data sets included in multiple non-faulty dedicated processors form updated optimizer data includes: determining multiple consecutive identifiers according to the identifiers included in the integrity information sent by non-faulty dedicated processors in the computing system, where the data indicated by the multiple consecutive identifiers forms updated optimizer data; determining that the data indicated by the multiple consecutive identifiers is correct according to the correct information included in the integrity information; and determining the non-faulty dedicated processors where the data indicated by the multiple consecutive identifiers is located as the multiple non-faulty dedicated processors that form the updated optimizer data.
[0019] In another possible implementation, the integrity information further includes the iteration number of the data in the data set included in the non-faulty dedicated processor; the method further includes: determining that the iteration numbers of the data indicated by the multiple consecutive identifiers are the same according to the iteration number included in the integrity information.
[0020] In this way, based on the data iteration number, it is ensured that complete optimizer data is formed based on the latest data.
[0021] In another possible implementation, after the general-purpose processor instructs to persist the data sets included in multiple non-faulty dedicated processors, the method further includes: when restarting model training, the general-purpose processor obtains updated optimizer data; the general-purpose processor configures the updated first data set for the first dedicated processor; the general-purpose processor configures the updated first data set for the third dedicated processor.
[0022] In another possible implementation, the method further includes: the first dedicated processor converts the optimizer parameters in the updated first data set into model parameters; the first dedicated processor transmits the model parameters to the dedicated processor in the computing system for training the model parameters.
[0023] Since the end-of-life checkpoint only saves the optimizer data and does not save the model parameters, when resuming model training, the model parameters are calculated based on the optimizer data in the storage system and then transmitted to the dedicated processor for training the model parameters, so as to ensure the consistency of the optimizer data and the model parameters in the end-of-life checkpoint, and overcome the problem that the optimizer data and the model parameters are inconsistent due to the failure of the AllGather operation when the dedicated processor fails. In addition, the method provided in this application can save the checkpoint at the time of failure and resume model training according to the checkpoint at the time of failure, thereby effectively reducing the training loss and improving the efficiency of model training.
[0024] In another possible implementation, the optimizer data includes optimizer parameters and optimizer states, and the optimizer states include variance and momentum.
[0025] In a second aspect, a data processing method is provided. The computing system to which the method is applied includes a general-purpose processor and multiple dedicated processors, and the multiple dedicated processors are used to train an artificial intelligence model. Among them, the multiple dedicated processors include at least two copies of optimizer data, and the optimizer data includes a first data set, and at least two dedicated processors both include the first data set. For example, the multiple dedicated processors include a first dedicated processor and a second dedicated processor, and both the first dedicated processor and the second dedicated processor include the first data set. The method is executed by the first dedicated processor, and the method includes: training the artificial intelligence model to obtain a first gradient; updating the first data set according to the first gradient to obtain an updated first data set; transmitting the first gradient to the second dedicated processor.
[0026] In a possible implementation, the method further includes: obtaining the initial value of the first data set configured by the general-purpose processor.
[0027] In another possible implementation, the method further includes: converting the optimizer parameters in the updated first data set into model parameters; transmitting the model parameters to the dedicated processor in the computing system for training the model parameters.
[0028] In another possible implementation, the method further includes: persisting the included data set.
[0029] In a third aspect, a data processing method is provided. The computing system to which the method is applied includes a general-purpose processor and multiple dedicated processors, and the multiple dedicated processors are used to train an artificial intelligence model. Among them, the multiple dedicated processors include at least two copies of optimizer data, the optimizer data includes a first data set, and at least two dedicated processors both include the first data set. For example, the multiple dedicated processors include a first dedicated processor and a second dedicated processor, and both the first dedicated processor and the second dedicated processor include the first data set. The method is executed by the second dedicated processor, and the method includes: obtaining a first gradient sent by the first dedicated processor; updating the first data set according to the first gradient to obtain an updated first data set.
[0030] In a possible implementation, the method further includes: obtaining an initial value of the first data set sent by the first dedicated processor.
[0031] In another possible implementation, the method further includes: converting the optimizer parameters in the updated first data set into model parameters; transmitting the model parameters to the dedicated processor in the computing system for training the model parameters.
[0032] In another possible implementation, the method further includes: persisting the included data set.
[0033] In another possible implementation, the method further includes: training the artificial intelligence model to obtain a second gradient.
[0034] In a fourth aspect, a data processing method is provided. The computing system to which the method is applied includes a general-purpose processor and multiple dedicated processors, and the multiple dedicated processors are used to train an artificial intelligence model. Among them, the multiple dedicated processors include at least two copies of optimizer data, the optimizer data includes a first data set, and at least two dedicated processors both include the first data set. For example, the multiple dedicated processors include a first dedicated processor and a second dedicated processor, and both the first dedicated processor and the second dedicated processor include the first data set. The method is executed by the general-purpose processor, and the method includes: when the second dedicated processor fails, instructing to persist the data sets included in multiple non-failed dedicated processors. The data sets included in the multiple non-failed dedicated processors form updated optimizer data. The multiple non-failed dedicated processors include the first dedicated processor, and the updated optimizer data includes an updated first data set.
[0035] In a possible implementation, the general-purpose processor instructs to persist the data sets included in multiple non-faulty dedicated processors, including: obtaining integrity information sent by the non-faulty dedicated processors in the computing system, where the integrity information is used to indicate the integrity of the data sets included in the non-faulty dedicated processors; determining, according to the integrity information sent by the non-faulty dedicated processors in the computing system, that the data sets included in the multiple non-faulty dedicated processors form updated optimizer data; and sending a persistence command to the multiple non-faulty dedicated processors, where the persistence command is used to instruct to persist the data sets included in the multiple non-faulty dedicated processors.
[0036] In another possible implementation, the integrity information includes the identifiers and correct information of the data in the data set; determining, according to the integrity information sent by the non-faulty dedicated processors in the computing system, that the data sets included in the multiple non-faulty dedicated processors form updated optimizer data includes: determining multiple consecutive identifiers according to the identifiers included in the integrity information sent by the non-faulty dedicated processors in the computing system, where the data indicated by the multiple consecutive identifiers forms the updated optimizer data; determining that the data indicated by the multiple consecutive identifiers is correct according to the correct information included in the integrity information; and determining the non-faulty dedicated processors where the data indicated by the multiple consecutive identifiers is located as the multiple non-faulty dedicated processors that form the updated optimizer data.
[0037] In another possible implementation, the integrity information further includes the iteration times of the data in the data sets included in the non-faulty dedicated processors; the method further includes: determining that the iteration times of the data indicated by the multiple consecutive identifiers are the same according to the iteration times included in the integrity information.
[0038] In another possible implementation, the method further includes: when restarting model training, obtaining the updated optimizer data; configuring the updated first data set to the first dedicated processor; and configuring the updated first data set to the third dedicated processor.
[0039] In a fifth aspect, a data processing device is provided, where the data processing device includes each module for executing the methods of the dedicated processors in the first aspect or any possible design of the first aspect. For example, the data processing device includes a communication module, a training module, and an optimizer data update module.
[0040] The training module is used to train an artificial intelligence model to obtain a first gradient. The optimizer data update module is used to update a first data set according to the first gradient to obtain an updated first data set.
[0041] In a possible implementation, the communication module is used to transmit the first gradient.
[0042] In another possible implementation, the communication module is further configured to obtain the initial value of the first data set configured by the general-purpose processor.
[0043] In another possible implementation, the optimizer data update module is further configured to convert the optimizer parameters in the updated first data set into model parameters. The communication module is further configured to transmit the model parameters to the dedicated processor in the computing system for training the model parameters.
[0044] In another possible implementation, the optimizer data update module is further configured to persist the included data set.
[0045] In a sixth aspect, a data processing apparatus is provided. The data processing apparatus includes each module of the dedicated processor for executing the method in the first aspect or any possible design of the first aspect. For example, the data processing apparatus includes a communication module, a training module, and an optimizer data update module.
[0046] The communication module is configured to obtain the first gradient; the optimizer data update module is configured to update the first data set according to the first gradient to obtain the updated first data set.
[0047] In a possible implementation, the communication module is further configured to obtain the initial value of the first data set.
[0048] In another possible implementation, the optimizer data update module is further configured to convert the optimizer parameters in the updated first data set into model parameters. The communication module is further configured to transmit the model parameters to the dedicated processor in the computing system for training the model parameters.
[0049] In another possible implementation, the optimizer data update module is further configured to persist the included data set.
[0050] In a seventh aspect, a data processing apparatus is provided. The data processing apparatus includes each module of the general-purpose processor for executing the method in the first aspect or any possible design of the first aspect. For example, the data processing apparatus includes a communication module, a fault handling module, and an optimizer data update module.
[0051] The fault handling module is configured to, when a dedicated processor in the system fails, instruct to persist the data sets included in multiple non-failed dedicated processors. The data sets included in the multiple non-failed dedicated processors form the updated optimizer data. The multiple non-failed dedicated processors include the first dedicated processor. The updated optimizer data includes the updated first data set.
[0052] In a possible implementation, a communication module is configured to obtain integrity information sent by non-faulty dedicated processors in a computing system, where the integrity information is used to indicate the integrity of the data sets included in the non-faulty dedicated processors;
[0053] When the fault handling module instructs to persist the data sets included in multiple non-faulty dedicated processors, it is specifically configured to: determine, according to the integrity information sent by non-faulty dedicated processors in the computing system, that the data sets included in the multiple non-faulty dedicated processors form updated optimizer data; the communication module is further configured to send a persistence command to the multiple non-faulty dedicated processors, where the persistence command is used to instruct to persist the data sets included in the multiple non-faulty dedicated processors.
[0054] In another possible implementation, the integrity information includes the identifiers and correct information of the data in the data set; when the fault handling module determines, according to the integrity information sent by non-faulty dedicated processors in the computing system, that the data sets included in the multiple non-faulty dedicated processors form updated optimizer data, it is specifically configured to: determine multiple consecutive identifiers according to the identifiers included in the integrity information sent by non-faulty dedicated processors in the computing system, where the data indicated by the multiple consecutive identifiers forms the updated optimizer data; determine that the data indicated by the multiple consecutive identifiers is correct according to the correct information included in the integrity information; and determine the non-faulty dedicated processors where the data indicated by the multiple consecutive identifiers is located as the multiple non-faulty dedicated processors that form the updated optimizer data.
[0055] In another possible implementation, the integrity information further includes the iteration times of the data in the data sets included in the non-faulty dedicated processors; the fault handling module is further configured to determine that the iteration times of the data indicated by the multiple consecutive identifiers are the same according to the iteration times included in the integrity information.
[0056] In another possible implementation, the communication module is further configured to obtain the updated optimizer data when restarting model training; the optimizer data update module is configured to configure the updated first data set to the first dedicated processor; the communication module is further configured to configure the updated first data set to the third dedicated processor.
[0057] In an eighth aspect, a computing system is provided, which includes a general-purpose processor and multiple dedicated processors, and a memory is used to store a set of computer instructions; when the processor executes the set of computer instructions, the general-purpose processor and the multiple dedicated processors jointly execute the operation steps of the method in the first aspect or any possible implementation manner of the first aspect.
[0058] In a ninth aspect, a computer device is provided. The computer device includes a memory and a plurality of processors. The memory is used to store a set of computer instructions. When the processors execute the set of computer instructions, the plurality of processors jointly execute the operation steps of the method in the first aspect or any possible implementation manner of the first aspect.
[0059] In a tenth aspect, a computer-readable storage medium is provided, including: computer software instructions. When the computer software instructions run in a processor, the processor is caused to execute the operation steps of the method described in the first aspect or any possible implementation manner of the first aspect.
[0060] In an eleventh aspect, a computer program product is provided. When the computer program product runs on a computer, the computer is caused to execute the operation steps of the method described in the first aspect or any possible implementation manner of the first aspect.
[0061] For the technical effects brought by any one of the design manners in the second aspect to the eleventh aspect, reference may be made to the technical effects brought by the first aspect or different design manners in the first aspect, which will not be elaborated herein.
[0062] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a schematic diagram of large model training provided by the present application;
[0064] Figure 2 It is a schematic diagram of a checkpoint provided by the present application;
[0065] Figure 3 It is a schematic diagram of the logical process of model training provided by the present application;
[0066] Figure 4 It is a schematic diagram of the architecture of a data processing system provided by the present application;
[0067] Figure 5 It is a schematic diagram of the flow of a data processing method provided by the present application;
[0068] Figure 6 It is a schematic diagram of the flow of a data processing method provided by the present application;
[0069] Figure 7 It is a schematic diagram of gradient transmission provided by the present application;
[0070] Figure 8 It is a schematic diagram of end-of-life checkpoint saving provided by the present application;
[0071] Figure 9A flowchart of a data processing method provided by this application;
[0072] Figure 10 A schematic diagram of optimizer data storage provided by this application;
[0073] Figure 11 A schematic diagram of the structure of a data processing device provided by this application;
[0074] Figure 12 A schematic diagram of the structure of a data processing device provided by this application;
[0075] Figure 13 A schematic diagram of the structure of a data processing device provided by this application;
[0076] Figure 14 A schematic diagram of the structure of a computer device provided by this application. Detailed implementation
[0077] For ease of understanding, the main terms involved in this application are first explained.
[0078] Large model: It refers to an extremely large-scale artificial intelligence (AI) model. Large models are widely used in the field of natural language processing and are revolutionizing the state of natural language processing (NLP) tasks, giving rise to more powerful and intelligent language technologies. Large models are an important direction of AI development. Large models also have the ability to perform well in various natural language processing tasks, such as text classification, sentiment analysis, summary generation, translation, etc. Large models can be used in multiple application fields such as automatic writing, chatbots, virtual assistants, voice assistants, automatic translation, etc. For example, large language models (LLMs), Bidirectional Encoder Representations from Transformers (BERT), Generative Pre-Trained Transformer (GPT), GPT3, GPT4, MoE, etc. Large models have the following characteristics.
[0079] 1. Gigantic scale: The model size can reach hundreds of gigabytes (GB) or even larger. Large models contain billions, hundreds of billions, or even trillions of model parameters. Such a gigantic-scale model provides powerful expressive and learning capabilities.
[0080] 2. Multi-task learning: Large models can handle a variety of different natural language processing (NLP) tasks, such as machine translation, text summarization, question answering systems, etc., enabling the model to learn a more extensive and generalized language understanding ability.
[0081] 3. Powerful computing resources: Training large models usually requires hundreds or even thousands of dedicated processors and a large amount of time. For example, training large models takes several weeks to several months. Powerful computing resources can accelerate the training process while retaining the capabilities of large models. Dedicated processors include, but are not limited to, graphic processing units (GPUs), data processing units (DPUs), neural processing units (NPUs), and neural-network processing units (NPUs).
[0082] 4. Abundant data: Use a large amount of training data to train large models and leverage the scale advantage of the model parameters of large models.
[0083] Model parameters: refer to the variables or weights that need to be learned or adjusted in an artificial intelligence model. Model parameters can affect the prediction ability and performance of the model. During the model training process, the model tries different parameter combinations to optimize the performance of the model. Common model parameters include weights, biases, learning rates, and regularization coefficients, etc. When using the model for prediction, these model parameters will be used to calculate the output results.
[0084] Model training: refers to training an artificial intelligence model using a training set so that the artificial intelligence model can predict or classify unknown data. During the model training process, learning is carried out based on the features and target values in the training set, and a model is generated after training is completed. This model can be used to predict or classify unknown data. Model training is one of the most important links in machine learning and directly affects the accuracy and reliability of the model.
[0085] Optimizer: It can refer to an algorithm used to adjust the model parameters of an artificial intelligence model so that the artificial intelligence model can more accurately predict the output results. The purpose of the optimizer is to minimize the loss function, that is, the difference between the predicted value and the actual value of the model. Common optimizers include stochastic gradient descent (SGD), adaptive moment estimation (Adam), adaptive gradient (Adagrad), and root mean square prop (RMSprop), etc. These optimizers use different strategies to update the model parameters to enable the artificial intelligence model to achieve better performance and accuracy. Usually, the input of the optimizer is the gradient, and the output is the model parameters.
[0086] Checkpoint (CKPT): It refers to saving checkpoint data at regular intervals or training epochs during the model training process for subsequent model restoration and continued training. The checkpoint includes model parameters and optimizer data. Optimizer data includes optimizer parameters and optimizer state (OS). The optimizer state includes variance and momentum. For example, if an unexpected situation occurs during the model training process resulting in the interruption of model training, the model training can be resumed based on the latest checkpoint without having to start the model training from scratch. Checkpoints can also be used for model evaluation and debugging. For example, observing the performance of the model at different training stages based on different checkpoints to judge the training effect and optimization direction of the model. In deep learning, checkpoints are a very important way to save and manage model parameters.
[0087] Resume training from breakpoint: It is a deep learning training technique that allows pausing the model training during the model training process and saving the current data. Generally, the current data includes model parameters and optimizer data. When restarting the training, the model can continue training from the place where it was paused last time without having to start from the beginning. This technique can accelerate the model training process, reduce the waste of computing resources, and can handle various problems that occur during the model training process, such as computer failures or network interruptions.
[0088] Figure 1 A schematic diagram of large model training provided for this application. As Figure 1As shown, the computing system 100 includes multiple cards. The cards can also be referred to as training cards, acceleration cards, or dedicated processors. The multiple cards can be located in multiple servers. The servers can be referred to as hosts. Parallel technologies are used based on the multiple cards to accelerate the training of large models. The parallel technologies include DataParallel, Model Parallel, Pipeline Parallel, and Optimizer Parallel.
[0089] DataParallel means dividing the multiple cards into multiple data parallel domains. As Figure 1 shown, the multiple cards are divided into X data parallel domains. The training set is sliced based on the data parallel domains, and the cards included in each data parallel domain train the large model based on the sub-training set. The sub-training sets used by the multiple data parallel domains to train the large model can be different.
[0090] Model Parallel means slicing the large model into multiple sub-models. For example, slicing the large model into multiple layers, and each card can train at least one layer of the model. Another example is that if the scale of a single layer of the model is large, it can also be sliced, and multiple cards can train one layer of the model. Different data parallel domains can train different layers of the large model. The multiple data parallel domains can train the large model in parallel.
[0091] Pipeline Parallel means slicing the large model into multiple layers according to the logical order between the multiple layers in the large model, and training the layers of the large model in parallel or serially by multiple cards according to the logical order between the multiple layers in the large model. The logical order between the layers of the large model can refer to the dependency relationship between the layers. For example, if the output data of the first layer is the input data of the second layer, then two cards can train the first layer and the second layer serially.
[0092] Optimizer Parallel means dividing the optimizer data into multiple data sets based on the number of cards in the computing system. Each card contains a part of the data in the optimizer data, thereby reducing the storage requirement of the optimizer data for the cards. The data sets saved by each card are different, and there is only one copy of a data set globally. The data sets saved by multiple cards can form the complete optimizer data.
[0093] For example, during LLM training, the data stored in the card mainly includes activations, model parameters, and optimizer data. The activations and model parameters are represented in half-precision floating-point (FP16), and the optimizer data is represented in single-precision floating-point (e.g., FP32). Assume the number of model parameters is M, the storage space required for model parameters is 2M Bytes, and the storage space required for optimizer data is 12M Bytes. That is, the optimizer data includes M optimizer parameters, M variances, and M momentums. The storage space required for M optimizer parameters is 4M Bytes. The storage space required for M variances is 4M Bytes. The storage space required for M momentums is 4M Bytes. For example, GPT3 includes 175 billion model parameters, and the storage space required for optimizer data is 12 bytes * 175 billion = 2.45 Terabytes (TB). For distributed training of LLM with multi-machine and multi-card cooperation, each card stores a part of the optimizer data, and multiple cards store the complete optimizer data.
[0094] Among them, each card performs operations such as forward, activation, backward, gradient generation, and gradient accumulation, and distributes the gradients to multiple cards through the AllReduce operation. The card updates the optimizer data with the same identifier as the gradient according to the gradient; obtains the model parameters converted from the optimizer parameters through the AllGather operation, thereby realizing model training.
[0095] In addition, there are multiple copies of model parameters globally, each data parallel domain contains a complete set of model parameters, and the model parameters included in the cards in different data parallel domains are the same. For example, data parallel domain 1 contains M model parameters, and data parallel domain 2 contains M model parameters. Among them, the M model parameters are distributed and stored on the cards in the data parallel domain. For example, card 1, card 5, and card x include parameter 1. Card 2, card 6, and card x + 1 include parameter 2.
[0096] There is one copy of optimizer data globally. Assume the system includes n cards, and each card includes 1 / n of the optimizer data. The n copies of 1 / n optimizer data included in the n cards form the complete optimizer data.
[0097] In some embodiments, the card can periodically save checkpoints, that is, save a checkpoint every once in a while. The checkpoint can be called a periodic checkpoint. After the model training fails at any time point, resume the model training according to the previous checkpoint.
[0098] Exemplarily, as shown in (a) of Figure 2 when card n fails, card 1 and card 2 persist the model parameters and optimizer data to the storage system. As shown in Figure 2As shown in (b) in [reference], during the model training process, the i-th checkpoint and the (i + 1)-th checkpoint are saved. After the (i + 1)-th checkpoint, the n-th card fails and the model training is interrupted. The fault is detected, the fault mode is judged, and whether to restart the training. If the model training is restarted, the model training is resumed according to the (i + 1)-th checkpoint.
[0099] Since the model training is resumed using the previous checkpoint, it will result in some training losses. For example, the updated data of the optimizer data and the model parameters are lost when training the model from the (i + 1)-th checkpoint to the fault point.
[0100] Measuring the training loss mainly includes two factors: 1. The number of cards used for the training task, that is, when training a large model, thousands of cards are used. Even for a loss of a few minutes, the number of cards * loss time will result in a relatively serious training loss. 2. The failure probability of the cards and the network, that is, when the failure rate of the cards or the network is relatively high, the frequency of training interruption will increase, resulting in an increase in the training loss.
[0101] The training loss can be used to calculate the efficiency of multiple cards. The more serious the training loss, the lower the efficiency of the cards.
[0102] The ways to reduce the training loss mainly include two aspects: 1. Reduce the failure probability; 2. Try to remedy through the last state when a failure occurs. For example, try to save the checkpoint at the time of failure and resume the model training according to the checkpoint at the time of failure to reduce the training loss. The checkpoint at the time of failure can be called the dying checkpoint.
[0103] There are also two aspects of problems in how to implement the technology of the dying checkpoint:
[0104] 1. Integrity of the checkpoint: As Figure 1 it can be seen, there is only one copy of the optimizer data globally, and each card contains 1 / n of the optimizer data. Once a card fails, it will lead to an incomplete checkpoint, which also leads to incomplete optimizer data.
[0105] 2. Data consistency of the checkpoint: Since the optimizer data is incomplete, the model parameters cannot be updated according to the optimizer data, resulting in inconsistent optimizer data and model parameters, and the model training converges slowly, or even worse than resuming the model training according to the previous periodic checkpoint.
[0106] The situation of inconsistent optimizer data and model parameters mainly includes that the optimizer data is the latest optimizer data, the model parameters are the model parameters of the previous iteration, or the model parameters on some cards are the latest model parameters, but the model parameters on another part of the cards are the model parameters of the previous step. The main reason for this situation is the failure of the AllGather operation. For example, Figure 3As shown in the figure, it is a schematic diagram of the logical process of model training provided by this application. The gradient is obtained through the AllReduce operation, the optimizer data is updated according to the gradient, and the optimizer data of checkpoint 1 is saved. The optimizer parameters are converted into model parameters. For example, the optimizer parameter fp32 is converted into the model parameter fp16. However, due to a card failure, the AllGather operation fails and the model parameters of checkpoint 2 cannot be updated, resulting in inconsistent optimizer data and model parameters. Among them, the data volume of the optimizer data is 12M. Taking M = 175B as an example, the size of the optimizer data is 2.4TB, and the size of the model parameters is 400GB.
[0107] To solve the problem that the final checkpoint is inaccurate and the optimizer data is incomplete when a dedicated processor fails, this application provides a data processing method. The computing system to which the method is applied includes a general-purpose processor and multiple dedicated processors, and the multiple dedicated processors are used to train an artificial intelligence model. Among them, the multiple dedicated processors include at least two copies of optimizer data, and the optimizer data includes a first data set, and at least two dedicated processors both include the first data set. For example, the multiple dedicated processors include a first dedicated processor and a second dedicated processor, and both the first dedicated processor and the second dedicated processor include the first data set. The method includes: the first dedicated processor trains the artificial intelligence model to obtain a first gradient; the first dedicated processor updates the first data set according to the first gradient to obtain an updated first data set; the second dedicated processor updates the first data set according to the first gradient to obtain an updated first data set. When the second dedicated processor fails, the general-purpose processor instructs to persist the data sets included in multiple non-failed dedicated processors, and the data sets included in the multiple non-failed dedicated processors form the updated optimizer data. The multiple non-failed dedicated processors include the first dedicated processor, and the updated optimizer data includes the updated first data set.
[0108] Compared with having one copy of optimizer data globally, when a dedicated processor fails, the final checkpoint is inaccurate and the optimizer data is incomplete. In the method provided by this application, at least two dedicated processors globally include the data sets in the optimizer data, that is, at least two copies of optimizer data are deployed globally to achieve multiple copies of optimizer data globally. Also, the dedicated processors update the data sets in the optimizer data they contain by transmitting gradients, that is, update multiple copies of optimizer data through "computing instead of transmitting", ensuring the integrity of the optimizer data at any time. When a dedicated processor fails, since the system includes multiple copies of optimizer data, the complete optimizer data can be composed of the data sets included in non-failed dedicated processors, improving the reliability of the optimizer data, ensuring the accuracy of the final checkpoint, facilitating the recovery of model training based on the final checkpoint, and reducing training losses.
[0109] The following describes in detail the implementation manners of the data processing method provided in this application with reference to the accompanying drawings.
[0110] Figure 4 It is a schematic architecture diagram of a data processing system provided in this application. As Figure 4 shown, the data processing system 400 includes a client 410, a computing cluster 420, and a storage cluster 430.
[0111] The storage cluster 430 includes multiple storage nodes 431. One storage node 431 includes one or more controllers, network cards, and multiple hard disks. The hard disks are used to store data. The hard disks can be magnetic disks or other types of storage media, such as solid-state drives or shingled magnetic recording hard disks, etc. The network cards are used to communicate with the computing nodes 421 included in the computing cluster 420. The controllers are used to write data to the hard disks or read data from the hard disks according to the read / write data requests sent by the computing nodes 421. During the process of reading and writing data, the controllers need to convert the addresses carried in the read / write data requests into addresses that the hard disks can recognize.
[0112] The computing cluster 420 includes multiple computing nodes 421. The computing nodes 421 can be a type of computing device, such as acceleration cards, servers, etc.
[0113] In some embodiments, the computing cluster 420 can be a heterogeneous computing architecture to provide high-performance computing. For example, the computing nodes 421 can include computing units with computing capabilities such as central processing units (CPUs), graphics processing units (GPUs), data processing units (DPUs), neural processing units (NPUs), and embedded neural network processors (NPUs) to provide high-performance computing.
[0114] In some other embodiments, multiple computing nodes 421 are connected based on high-speed interconnect technology through network devices (such as switches, network cards, etc.) to enable communication between multiple computing nodes 421.
[0115] The client 410 communicates with the computing cluster 420 and the storage cluster 430 via the network 440. For example, the client 410 sends a request to the computing cluster 420 via the network 440, requesting the computing cluster 420 to perform model training. The network 440 may refer to an enterprise internal network (such as a Local Area Network (LAN)) or the Internet. The client 410 refers to a computer connected to the network 440, and can also be called a workstation. Different clients can share resources on the network (such as computing resources and storage resources).
[0116] In some embodiments, the computing cluster 420 further includes a control node 422. For example, the control node and the computing nodes can be independent physical devices. Also, for example, the control node and multiple computing nodes can be located on the same physical device. The control node can be a CPU. The multiple computing nodes include computing units such as GPUs, NPUs, and DPUs. The control node 422 is used to manage and allocate tasks, and multiple tasks are executed in parallel by multiple computing nodes to improve the data processing rate.
[0117] In this application, the control node 422 is used to instruct multiple computing nodes to perform model training based on the parallel technology described in the above embodiments according to the request.
[0118] The control node 422 is further used to divide multiple computing nodes in the system into multiple groups, each group contains at least two computing nodes, and each group contains a complete copy of the optimizer data, realizing global multi-copy optimizer data.
[0119] The control node 422 is further used to instruct non-faulty computing nodes to persistently store the optimizer data when a computing node fails.
[0120] Gradients can be transmitted between computing nodes 421 in different groups, so that the computing nodes 421 can update the data set in the optimizer data according to the gradients.
[0121] In the embodiments of this application, the storage cluster 430 stores the optimizer data and so on.
[0122] In some other embodiments, the client 410 installs a client program 411. The client 410 runs the client program 411 to display a user interface (UI), and the user 450 operates the user interface to submit a request. For example, the user 450 operates the user interface to submit a request. After the control node 422 obtains the request, it loads the optimizer data from the storage cluster 430, and distributes the optimizer data to multiple computing nodes, so that the multiple computing nodes perform model training based on the parallel technology described in the above embodiments.
[0123] Optionally, the system administrator 460 can configure system information, etc. by invoking the application platform interface (API) 412 or the command-line interface (CLI) interface 413 through the client 410. For example, the initial values of the optimizer data configured for the computing nodes provided in this application.
[0124] Figure 4 It is only a schematic diagram. The embodiments of this application do not limit the device connection method and the number of devices in the data processing system. For example, the data processing system may include multiple clients. One client can be connected to multiple computing nodes. Different clients establish connections with different computing nodes.
[0125] Next, the data processing process will be described in detail with reference to the accompanying drawings.
[0126] Figure 5 It is a flowchart of a data processing method provided in this application. Here, the optimizer data update and the end-of-life checkpoint are mainly described. Assume that the computing system includes a general-purpose processor and K dedicated processors.
[0127] As Figure 5 shown in (a) of, the initialization process before the dedicated processor trains the model. The method includes the following steps 510 to step 530.
[0128] Step 510, the general-purpose processor obtains a request.
[0129] The general-purpose processor obtains a request sent by the client. The request is used to indicate the execution of model training. The request may include a model identifier, so that the general-purpose processor can identify the model according to the model identifier and instruct the dedicated processor to load the model indicated by the model identifier from the storage system.
[0130] The general-purpose processor can obtain optimizer data from the storage system. It should be noted that when the general-purpose processor first instructs the dedicated processor to train the model, the general-purpose processor can obtain the initial optimizer data from the storage system. The initial optimizer data includes the initial values of the optimizer parameters, the initial values of the momentum, and the initial values of the variance. When the general-purpose processor instructs the dedicated processor to resume model training, that is, when the dedicated processor performs breakpoint continuation training, the general-purpose processor can obtain the updated optimizer data from the storage system. The updated optimizer data includes the updated values of the optimizer parameters, the updated values of the momentum, and the updated values of the variance.
[0131] Step 520, the general-purpose processor initializes the optimizer data.
[0132] The general-purpose processor can group the dedicated processors in the system to obtain multiple groups. Each group contains at least two dedicated processors. The number of dedicated processors in each group can be the same or different.
[0133] The general-purpose processor divides the optimizer data according to the number of dedicated processors in the group, and each dedicated processor in the group contains a part of the data in the optimizer data. At least two dedicated processors belonging to different groups contain the same optimizer data. Each group contains a complete copy of the optimizer data, and multiple groups contain multiple complete copies of the optimizer data, realizing global multi-copy optimizer data.
[0134] Exemplarily, the general-purpose processor divides K dedicated processors in the system into R groups. Each group contains at least two dedicated processors. The value of R is an integer greater than or equal to 2. The larger the number of groups R, the more copies of the optimizer data in the system, and the easier it is to obtain the complete optimizer data, improving the reliability of the optimizer data.
[0135] For ease of description, the following takes each group containing N dedicated processors as an example.
[0136] The general-purpose processor selects one of the R groups, and according to the number N of dedicated processors in the group, divides the optimizer data into N parts. The N dedicated processors in the group contain N parts of data, that is, one dedicated processor in the group includes 1 / N of the data in the optimizer data. The N parts of data contained by the N dedicated processors form the complete optimizer data. The value of N is an integer greater than or equal to 2.
[0137] In some embodiments, the optimizer data includes M optimizer parameters, M variances, and M momenta. The general-purpose processor can divide the M optimizer parameters, M variances, and M momenta into N parts. Alternatively described, the general-purpose processor can divide the M optimizer parameters, M variances, and M momenta into N data sets. One dedicated processor in the group contains one data set. Each data set contains optimizer parameters, variances, and momenta. The data sets contained by different dedicated processors in the group are different. The N data sets contained by the N dedicated processors form the complete optimizer data. The general-purpose processor configures the first data set to dedicated processor 1 of group 1, the second data set to dedicated processor 2 of group 1, and so on, and configures the Nth data set to dedicated processor N of group 1.
[0138] Among them, the number of optimizer parameters, the number of variances, and the number of momenta contained in the data set are the same. The data set contains one or more groups of optimizer parameters, variances, and momenta. For example, the data set contains one optimizer parameter, one variance, and one momentum. Another example is that the data set contains multiple optimizer parameters, multiple variances, and multiple momenta.
[0139] Different data sets may contain the same number of optimizer parameters. For example, each data set contains M / N optimizer parameters, M / N variances, and M / N momentums.
[0140] Different data sets may also contain different numbers of optimizer parameters. For example, the first data set contains 100 optimizer parameters, 100 variances, and 100 momentums. The second data set contains 150 optimizer parameters, 150 variances, and 150 momentums.
[0141] If different data sets contain the same number of optimizer parameters, the number of momentums and variances contained in different data sets is also the same. If different data sets contain different numbers of optimizer parameters, the number of momentums and variances contained in different data sets is also different.
[0142] Exemplarily, there are 2 dedicated processors in the group, and the optimizer data is divided into 2 data sets. The 2 dedicated processors in the group include 2 data sets, that is, the first dedicated processor in the group includes the first data set in the optimizer data, and the second dedicated processor in the group includes the second data set in the optimizer data. The 2 data sets contained in the 2 dedicated processors form the complete initial optimizer data. The optimizer data includes 175 billion optimizer parameters, 175 billion variances, and 175 billion momentums. Each dedicated processor in the group includes 87.5 billion optimizer parameters, 87.5 billion variances, and 87.5 billion momentums. The first dedicated processor contains the optimizer parameters, variances, and momentums identified from 1 to 87.5 billion. The first dedicated processor contains the optimizer parameters, variances, and momentums identified from 87.6 billion to 175 billion.
[0143] Step 530: The dedicated processor initializes the optimizer data.
[0144] The dedicated processors in the group transfer the configured optimizer data to the dedicated processors in other groups, so that the dedicated processors in other groups contain the optimizer data, so that each group contains a complete copy of the optimizer data, and multiple complete copies of the optimizer data contained in multiple groups realize the global multi-copy optimizer data.
[0145] For example, the dedicated processor 1 of group 1 transfers the first data set to the dedicated processor 1 of group R, the dedicated processor 2 of group 1 transfers the second data set to the dedicated processor 2 of group R, and so on. The dedicated processor N of group 1 transfers the Nth data set to the dedicated processor N of group R.
[0146] Exemplarily, assume R equals 2 and N equals 2, as Figure 6As shown in (a) of , the dedicated processors in the system are divided into 2 groups, with each group containing 2 dedicated processors. The optimizer data includes a first data set and a second data set. The dedicated processor 1 in group 1 contains the first data set, and the dedicated processor 2 in group 1 contains the second data set. The dedicated processor 1 in group 1 transmits the first data set to the dedicated processor 1 in group 2. The dedicated processor 2 in group 1 transmits the second data set to the dedicated processor 2 in group 2.
[0147] After the initialization process is completed, the dedicated processors train the model. As Figure 5 shown in (b) of , the process of updating the optimizer data. The method includes the following steps 540 to step 560.
[0148] Step 540: The dedicated processors train the model to obtain gradients.
[0149] The dedicated processors within each group train the model based on the training data to obtain gradients. For example, the dedicated processors perform operations such as forward, activation, backward, gradient generation, and gradient accumulation.
[0150] In some embodiments, the optimizer data includes M optimizer parameters, M variances, and M momentums. The dedicated processors within each data parallel domain train the model based on the training data to obtain M gradients.
[0151] Step 550: The dedicated processors transmit the gradients to the dedicated processors that contain the same data set.
[0152] The dedicated processors transmit the gradients to the dedicated processors where the optimizer data with the same identifier as the gradients is located.
[0153] For example, the dedicated processor 1 in group 1 and the dedicated processor 1 in group 2 both contain the first data set, and the first data set includes optimizer data with identifiers from 1 to M / 2. The dedicated processor 2 in group 1 and the dedicated processor 2 in group 2 both include the second data set, and the second data set includes optimizer data with identifiers from (M / 2)+1 to M.
[0154] The dedicated processor 1 in group 1 transmits the gradients with identifiers from 1 to M / 2 to the dedicated processor 1 in group 2, and the dedicated processor 2 in group 1 transmits the gradients with identifiers from (M / 2)+1 to M to the dedicated processor 2 in group 2.
[0155] Exemplarily, as Figure 7As shown, all the cards in the computing system are divided into 4 data parallel domains, and each data parallel domain contains 250 cards. The optimizer data includes M optimizer parameters, M variances, and M momentums. All the cards in the computing system are divided into 2 groups, and the optimizer data is divided into 250 data sets. One card in each group includes one data set. Each data set includes M / 250 optimizer parameters, M / 250 variances, and M / 250 momentums. The gradients generated by the cards in the 4 data parallel domains are transmitted to the cards in group 1 and group 2 where the optimizer data with the same identifier as the gradients is located.
[0156] Exemplarily, as Figure 6 shown in (b) of, the dedicated processor 1 in group 1 generates gradient 1 and gradient 2, and the dedicated processor 2 in group 1 generates gradient 3 and gradient 4. The dedicated processor 1 in group 1 transmits gradient 1 and gradient 2 to the dedicated processor 1 in group 2. The dedicated processor 2 in group 1 transmits gradient 3 and gradient 4 to the dedicated processor 1 in group 1 and the dedicated processor 1 in group 2.
[0157] The dedicated processor 1 in group 2 generates gradient 5 and gradient 6, and the dedicated processor 2 in group 2 generates gradient 7 and gradient 8. The dedicated processor 1 in group 2 transmits gradient 5 and gradient 6 to the dedicated processor 2 in group 1 and the dedicated processor 2 in group 2. The dedicated processor 2 in group 2 transmits gradient 7 and gradient 8 to the dedicated processor 2 in group 1.
[0158] In some embodiments, if the first data set contains the optimizer parameters, variances, and momentums corresponding to multiple identifiers, then the gradients corresponding to the multiple identifiers are transmitted to the dedicated processor where the data set with the same multiple identifiers is located. For example, if the first data set contains the optimizer parameters, variances, and momentums corresponding to identifiers 1 to 3, and both the dedicated processor 1 in group 1 and the dedicated processor 1 in group 2 contain the first data set, then the dedicated processor 1 in group 1 transmits the gradients corresponding to identifiers 1 to 3 to the dedicated processor 1 in group 2.
[0159] Step 560: The dedicated processor updates the optimizer data according to the gradients.
[0160] Each dedicated processor updates the data set in the optimizer data contained in the dedicated processor according to the gradients, and obtains the updated data set.
[0161] The dedicated processor 1 in group 1 and the dedicated processor 1 in group R update the first data set according to the gradients. And so on, the dedicated processor N in group 1 and the dedicated processor N in group R update the Nth data set according to the gradients.
[0162] For example, the dedicated processor 1 in group 1 and the dedicated processor 1 in group 2 both update the optimizer parameters, variance, and momentum (parameter momentum variance, PMV) of identifier 1 according to the gradient of identifier 1, and so on. The dedicated processor 1 in group 1 and the dedicated processor 1 in group 2 both update the optimizer parameters, variance, and momentum of identifier M / 2 according to the gradient of identifier M / 2, obtaining the updated optimizer parameters, updated variance, and updated momentum.
[0163] The dedicated processor 2 in group 1 and the dedicated processor 2 in group 2 both update the optimizer parameters, variance, and momentum of identifier (M / 2)+1 according to the gradient of identifier (M / 2)+1, and so on. The dedicated processor 2 in group 1 and the dedicated processor 2 in group 2 both update the optimizer parameters, variance, and momentum of identifier M according to the gradient of identifier M, obtaining the updated optimizer parameters, updated variance, and updated momentum.
[0164] The method for updating the optimizer parameters, variance, and momentum can refer to traditional methods and is not limited.
[0165] The optimizer data update scheme provided in this application includes ensuring that the optimizer data is the same among multiple groups during initialization. During the model training process, gradients are transmitted among multiple groups. Then, based on the same optimizer data (previous PMV), the same input (gradient), and the same optimizer algorithm, the dedicated processors make the updated optimizer data the same among multiple groups, that is, the update results (current PMV) are the same. Thus, since the data volume of the gradient is much smaller than the data volume of the optimizer data, for example, the data volume of the gradient is 1 / 6 of the data volume of the optimizer data, only the gradients are transmitted among the dedicated processors, reducing the transmitted data volume and bandwidth. The dedicated processors update the optimizer data according to the gradients, ensuring that the updated optimizer data is the same among multiple groups, achieving "computing instead of transmitting", and ensuring the integrity of the optimizer data at any time.
[0166] After the optimizer data update is completed, the dedicated processor updates the model parameters. As Figure 5 shown in (c) of
[0167] Step 570: The dedicated processor converts the optimizer parameters into model parameters.
[0168] The dedicated processor converts the single-precision optimizer parameters into half-precision model parameters. For example, converting the optimizer parameter fp32 into the model parameter fp16.
[0169] Step 580: The dedicated processor transmits the model parameters to the dedicated processor for training the model parameters.
[0170] After the dedicated processor converts the single-precision optimizer parameters into half-precision model parameters, it performs the AllGather operation. For example, each dedicated processor in R groups contains the model parameter P1. The dedicated processor 1 in group 1 transmits the model parameter P1 to each dedicated processor in the R groups. Each dedicated processor in the R groups contains the model parameter P2. The dedicated processor 2 in group 1 transmits the model parameter P2 to each dedicated processor in the R groups.
[0171] Exemplarily, as Figure 6 shown in (c) of, the dedicated processor 1 and the dedicated processor 2 in group 1, and the dedicated processor 1 and the dedicated processor 2 in group 2 both contain the model parameter P1 and the model parameter P2. The dedicated processor 1 in group 1 converts the single-precision optimizer parameter into the half-precision model parameter P1, and transmits the model parameter P1 to the dedicated processor 1 and the dedicated processor 2 in group 1, and the dedicated processor 1 and the dedicated processor 2 in group 2. The dedicated processor 2 in group 1 converts the single-precision optimizer parameter into the half-precision model parameter P2, and transmits the model parameter P2 to the dedicated processor 1 and the dedicated processor 2 in group 1, and the dedicated processor 1 and the dedicated processor 2 in group 2.
[0172] Thus, it ensures the consistency of the optimizer data and the model parameters, and ensures the correct model training.
[0173] In some other embodiments, after a dedicated processor in the system fails, the non-failed dedicated processors can persist the dying checkpoint without saving the model parameters. As Figure 8 shown, this application further includes the following steps 590 to step 5130.
[0174] Here, it is assumed that the dedicated processor 1 in group 1 fails, and the non-failed dedicated processors sense that there is an abnormal dedicated processor in the computing system. The general processor instructs to persist the data sets contained in multiple non-failed dedicated processors, and the data sets contained in multiple non-failed dedicated processors constitute the updated optimizer data. Multiple non-failed dedicated processors include the first dedicated processor, and the updated optimizer data includes the updated first data set. The data sets contained in multiple non-failed dedicated processors include the updated first data set.
[0175] Step 590: The non-failed dedicated processors send an exception message to the general processor.
[0176] If a dedicated processor in the computing system fails, other non-failed dedicated processors may not be able to receive the data sent by the failed dedicated processor, then the non-failed dedicated processors sense that there is an abnormal dedicated processor in the computing system and send an exception message to the general processor. The exception message is used to indicate that there is an abnormal dedicated processor in the system.
[0177] Step 5100, the general - purpose processor instructs the non - faulty dedicated processors to report integrity information.
[0178] The general - purpose processor sends a reporting command to the non - faulty dedicated processors, and the non - faulty dedicated processors report integrity information to the general - purpose processor. The reporting command is used to instruct the non - faulty dedicated processors to report integrity information.
[0179] The integrity information is used to indicate the integrity of the data set contained in the non - faulty dedicated processors. For example, the integrity information includes the identifiers and correct information of the data in the data set. The integrity information also includes the number of iterations of the data in the data set contained in the non - faulty dedicated processors.
[0180] Step 5110, determine that the data sets contained in multiple non - faulty dedicated processors in the computing system form updated optimizer data according to the integrity information sent by the non - faulty dedicated processors.
[0181] Since at least two copies of the data set contained in the optimizer data are stored in the computing system, the integrity information received by the general - purpose processor may include the same integrity information, that is, at least two dedicated processors send the same identifier to the general - purpose processor. It should be noted that if two copies of the data set in the optimizer data are stored in the system and one dedicated processor fails, the general - purpose processor receives the integrity information of the data set contained in the faulty dedicated processor from the non - faulty dedicated processors. The general - purpose processor can also receive the integrity information of the data sets contained in two non - faulty dedicated processors from other non - faulty dedicated processors.
[0182] The general - purpose processor filters multiple consecutive identifiers from the received identifiers, and the data indicated by the multiple consecutive identifiers forms the updated optimizer data.
[0183] The general - purpose processor determines the non - faulty dedicated processors where the data indicated by the multiple consecutive identifiers is located as the multiple non - faulty dedicated processors that form the updated optimizer data. The updated optimizer data includes updated optimizer parameters, updated variances, and updated momenta.
[0184] Understandably, the data contained in multiple non - faulty dedicated processors forms complete optimizer data. That is, the general - purpose processor determines the dedicated processors that can form complete optimizer data according to the multiple consecutive identifiers contained in the integrity information.
[0185] Optionally, the general - purpose processor can also determine that the data indicated by the multiple consecutive identifiers is correct according to the correct information contained in the integrity information, and then determine the non - faulty dedicated processors where the data indicated by the multiple consecutive identifiers is located as the multiple non - faulty dedicated processors that form the updated optimizer data.
[0186] Optionally, the integrity information further includes the number of iterations. If the general-purpose processor determines that the number of data iterations indicated by multiple consecutive identifiers is the same according to the number of iterations included in the integrity information, the non-faulty dedicated processors where the data indicated by the multiple consecutive identifiers are located are determined as the multiple non-faulty dedicated processors that make up the updated optimizer data.
[0187] Understandably, the multiple identifiers determined by the general-purpose processor for composing the complete optimizer data are consecutive, the number of iterations of the data indicated by the multiple identifiers is the same, and the data is all correct.
[0188] Exemplarily, the optimizer data includes a first data set and a second data set. Dedicated processor 1 in group 1 and dedicated processor 1 in group 2 contain the first data set. Dedicated processor 2 in group 1 and dedicated processor 2 in group 2 contain the second data set. Dedicated processor 1 in group 1 fails, and dedicated processor 1 in group 2 reports the integrity information of the first data set to the general-purpose processor. Dedicated processor 2 in group 1 and dedicated processor 2 in group 2 report the integrity information of the second data set to the general-purpose processor. The general-purpose processor determines that the first data set contained in dedicated processor 1 in group 2 and the second data set contained in dedicated processor 2 in group 1 compose the complete optimizer data. Then the general-purpose processor instructs dedicated processor 1 in group 2 to persist the first data set, and instructs dedicated processor 2 in group 1 to persist the second data set.
[0189] Step 5120: The general-purpose processor sends a persistence command to multiple non-faulty dedicated processors.
[0190] The persistence command is used to instruct to persist the data sets contained in multiple non-faulty dedicated processors. The general-purpose processor instructs multiple non-faulty dedicated processors to persist the data sets in the optimizer data they contain.
[0191] Step 5130: The dedicated processor persists the data.
[0192] Upon receiving the persistence command sent by the general-purpose processor, the dedicated processor persists the data sets in the optimizer data it contains to the storage system. For example, if the optimizer data includes a first data set and a second data set, the first data set in the optimizer data contained in dedicated processor 1 is stored in the storage system, and the second data set in the optimizer data contained in dedicated processor 2 is stored in the storage system. The storage system can store the optimizer data in the form of files.
[0193] Understandably, a dying checkpoint is stored in the storage system, that is, the complete optimizer data in the system when the dedicated processor fails. The optimizer data includes updated optimizer parameters, updated variances, and updated momenta.
[0194] Thus, due to the optimizer data with multiple copies in the global scope, when a dedicated processor fails, since the system contains multiple copies of the optimizer data, the complete optimizer data can be composed of the data sets contained in the non-failed dedicated processors, ensuring the integrity of the optimizer data when the dedicated processor fails, improving the reliability of the optimizer data, ensuring the accuracy of the dying checkpoint, facilitating the restoration of model training based on the dying checkpoint, and reducing the training loss.
[0195] Optionally, the above embodiments illustrate the execution of steps 5100 to 5130 by a general-purpose processor. In some embodiments, a dedicated processor can also be specified by the general-purpose processor to execute steps 5100 to 5130.
[0196] In some other embodiments, after the dedicated processor in the system fails, the dedicated processor can restore model training according to the dying checkpoint. Figure 9 It is a schematic flowchart of a data processing method provided by this application. Here, the restoration of model training is mainly described. As Figure 9 shown in (a) of [reference], in the initialization process before the dedicated processor trains the model, the method includes the following steps 910 to 930. After the optimizer data update is completed, the dedicated processor updates the model parameters. As Figure 9 shown in (b) of [reference], in the process of updating the model parameters. The method includes the following steps 940 to 960.
[0197] Step 910: The general-purpose processor obtains a request.
[0198] The general-purpose processor obtains the request sent by the client, and the request is used to indicate the restoration of model training. The request may include a model identifier, so that the general-purpose processor can identify the model according to the model identifier and instruct the dedicated processor to load the model indicated by the model identifier from the storage system.
[0199] The general-purpose processor can obtain the optimizer data from the storage system. Since the general-purpose processor instructs the dedicated processor to restore the training model, that is, when the dedicated processor performs breakpoint continuation training, the general-purpose processor obtains the updated optimizer data from the storage system. The updated optimizer data includes the updated values of the optimizer parameters, the updated values of the momentum, and the updated values of the variance. The updated optimizer data can be the optimizer data contained in the dying checkpoint.
[0200] Step 920: The general-purpose processor initializes the updated optimizer data.
[0201] The general-purpose processor can group the dedicated processors in the system to obtain multiple groups. The general-purpose processor selects one group from the multiple groups and divides the updated optimizer data according to the number of dedicated processors in the group. Each dedicated processor in the group contains a part of the data in the updated optimizer data. At least two dedicated processors belonging to different groups contain the same updated optimizer data. Each group contains a complete copy of the updated optimizer data, and the multiple complete copies of the updated optimizer data contained in the multiple groups implement the globally replicated updated optimizer data. For the explanation of how the general-purpose processor initializes the dedicated processors, reference can be made to the description of step 520 above.
[0202] It should be noted that after this grouping, the dedicated processors included in each group can be the same as or different from those included in each group after the previous grouping, without limitation. After the faulty dedicated processor is restored, this grouping can also include the faulty recovered dedicated processor.
[0203] Step 930: The dedicated processor initializes the updated optimizer data.
[0204] The dedicated processors within the group transmit the configured optimizer data to the dedicated processors in other groups, so that the dedicated processors in other groups contain the updated optimizer data, enabling each group to contain a complete copy of the updated optimizer data, and the multiple complete copies of the updated optimizer data contained in the multiple groups implement the globally replicated optimizer data.
[0205] It should be noted that step 930 is an optional step. In some embodiments, the general-purpose processor can also obtain the updated optimizer data from the storage system again, divide the updated optimizer data into N parts according to the number N of dedicated processors in another group, and configure 1 / N of the data in the updated optimizer data for one dedicated processor in the other group.
[0206] Step 940: The dedicated processor converts the optimizer parameters into model parameters.
[0207] The dedicated processor converts the single-precision optimizer parameters into half-precision model parameters. For example, converting the optimizer parameter fp32 into the model parameter fp16.
[0208] Step 950: The dedicated processor transmits the model parameters to the dedicated processor for training the model parameters.
[0209] The dedicated processor transmits the gradient to the dedicated processor where the optimizer data with the same identifier as the gradient is located. For the explanation of step 950, reference can be made to the description of step 580 above.
[0210] Step 960: The dedicated processor trains the model.
[0211] Since the end-of-life checkpoint only saves the optimizer data and does not save the model parameters, when resuming model training, the model parameters are calculated based on the optimizer data in the storage system, and then the model parameters are transmitted to the dedicated processor for training the model parameters, so as to ensure the consistency between the optimizer data and the model parameters of the end-of-life checkpoint, and overcome the problem that when the dedicated processor fails, due to the failure of the AllGather operation, the optimizer data and the model parameters are inconsistent. In addition, the method provided in this application can save the checkpoint at the time of failure and resume model training according to the checkpoint at the time of failure, thereby effectively reducing the training loss and improving the efficiency of model training.
[0212] Figure 10 It is a schematic diagram of the storage and comparison of optimizer data provided in this application. As Figure 10 shown in (a) therein, all the cards in the computing system are divided into 4 data parallel domains, and each data parallel domain contains 4 cards. The optimizer data includes M optimizer parameters, M variances, and M momentums. The optimizer data is divided into 16 parts, and each card contains one 1 / 16 data.
[0213] As Figure 10 shown in (b) therein, all the cards in the computing system are divided into 2 groups, the optimizer data is divided into 8 data sets, and one card in each group includes one data set. Each data set includes M / 8 optimizer parameters, M / 8 variances, and M / 8 momentums. The gradients generated by the cards in the 4 data parallel domains are transmitted to the cards in group 1 and group 2 where the optimizer data with the same identifier as the gradient is located.
[0214] Therefore, if the optimizer data is replicated violently, the data transmission volume is 12M, which is the data volume of the entire optimizer data. The method provided in this application has a data transmission volume of 2M, which is the data volume of the gradient. At least two copies of the optimizer data are deployed globally to implement multiple copies of the optimizer data globally.
[0215] Figure 11 It is a schematic diagram of a data processing device 1100 provided in this application. Among them, during the model training process, the optimizer multi-copy device 1101 is used to perform calculation instead of transmission to implement the update of multiple copies of the optimizer data and ensure the integrity of the optimizer data at any time.
[0216] When an exception occurs in the system, the save process of the end-of-life checkpoint includes a training exception acquisition device 1102, an optimizer data global save device 1103, and a distributed job cluster 1104. The training exception acquisition device 1102 is used to sense system exceptions. The optimizer data global save device 1103 is used to store the data sets in the optimizer data.
[0217] The working cluster 1104 is used to perform the health status detection of workers, execute the integrity check, and issue the persistent command of the end-of-life checkpoint.
[0218] Figure 12 It is a schematic diagram of another data processing device provided by this application. The data processing device 1200 includes a recovery device for the end-of-life checkpoint. The recovery device for the end-of-life checkpoint includes an optimizer state loading module 1201, a parameter generator 1202, and a parameter propagator 1203.
[0219] The optimizer state loading module 1201 is used to read the optimizer data from the end-of-life checkpoint file and synchronize the copy of the optimizer data.
[0220] The parameter generator 1202 is used to generate model parameters according to the optimizer parameters. Since the optimizer data contains optimizer parameters, but the types of optimizer parameters and model parameters are inconsistent. For example, the optimizer parameters are fp32, while the model parameters fp16 are used in the forward and backward operation processes. If the optimizer parameters and model parameters are inconsistent, then convert fp32 to fp16.
[0221] The parameter propagator 1203 is used to synchronize the model parameters to all corresponding training workers through communication operations. When the propagation of the model parameters is completed, the training iteration can start.
[0222] Figure 11 and Figure 12 For the explanation of the functions of each device in the data processing device shown, reference can be made to the description in the above embodiments.
[0223] It can be understood that in order to implement the functions in the above embodiments, the computer device includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combined with the units and method steps of each example described in the embodiments disclosed in this application, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application scenario and design constraint conditions of the technical solution.
[0224] In the above text, in combination with Figures 1 to 13 , the data processing method provided by this application is described in detail. Next, in combination with Figure 13 , the devices provided by this application will be described. These devices can be used to implement the functions of the dedicated processor or general-purpose processor in the above method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments. In this embodiment, the device can be such as Figure 5 , Figure 8 , Figure 9The dedicated processor or general - purpose processor shown may also be a module (such as a chip) applied to a computer device.
[0225] As Figure 13 shown in (a) of Figure 13 , the data - processing device 1300 includes a communication module 1301, a training module 1302, an optimizer data - update module 1303, and a storage module 1304.
[0226] The data - processing device 1300 is used to implement the functions of the dedicated processor in the method embodiments shown above Figure 5 , Figure 8 , Figure 9 .
[0227] The training module 1302 is used to train an artificial - intelligence model to obtain a first gradient. For example, the training module 1302 is used to execute step 960 in Figure 9 , and step 540 in Figure 5 .
[0228] The optimizer data - update module 1303 is used to update the first data set according to the first gradient to obtain an updated first data set. For example, the optimizer data - update module 1303 is used to execute step 560 in Figure 5 .
[0229] Optionally, the communication module 1301 is used to transmit the first gradient. For example, the communication module 1301 is used to execute step 550 in Figure 5 .
[0230] Optionally, the communication module 1301 is also used to obtain the initial value of the first data set configured by the general - purpose processor. For example, the communication module 1301 is used to execute step 510 in Figure 5 .
[0231] Optionally, the optimizer data - update module 1303 is also used to convert the optimizer parameters in the updated first data set into model parameters. For example, the optimizer data - update module 1303 is used to execute step 570 in Figure 5 , and step 940 in Figure 9 .
[0232] Optionally, the communication module 1301 is also used to transmit the model parameters to the dedicated processor in the computing system for training the model parameters. For example, the communication module 1301 is used to execute step 580 in Figure 5 , and step 950 in Figure 9 .
[0233] Optionally, the optimizer data - update module 1303 is also used to persist the included data set. For example, the optimizer data - update module 1303 is used to execute step 5130 in Figure 8 .
[0234] Optionally, the communication module 1301 is further configured to obtain the initial value of the first data set. For example, the communication module 1301 is configured to execute Figure 5 step 530 in
[0235] The storage module 1304 is used to store data sets, gradients, optimization algorithms, etc. in the optimizer data to facilitate model training.
[0236] As Figure 13 shown in (b) of Figure 8 the data processing device 1300 is used to implement the functions of the general-purpose processor in the method embodiments shown above in
[0237] The data processing device 1300 includes a communication module 1301, a fault handling module 1305, an optimizer data update module 1303, and a storage module 1304.
[0238] The fault handling module 1305 is configured to, when a dedicated processor in the system fails, instruct to persist the data sets included in multiple non-faulty dedicated processors. The data sets included in the multiple non-faulty dedicated processors form the updated optimizer data. The multiple non-faulty dedicated processors include a first dedicated processor, and the updated optimizer data includes an updated first data set. For example, the fault handling module 1305 is configured to execute Figure 8 step 5110 in
[0239] Optionally, the communication module 1301 is configured to obtain the integrity information sent by non-faulty dedicated processors in the computing system. The integrity information is used to indicate the integrity of the data sets included in the non-faulty dedicated processors. For example, the communication module 1301 is configured to execute Figure 8 step 5100 in
[0240] Optionally, when the fault handling module 1305 instructs to persist the data sets included in multiple non-faulty dedicated processors, it is specifically configured to: determine the data sets included in the multiple non-faulty dedicated processors to form the updated optimizer data according to the integrity information sent by non-faulty dedicated processors in the computing system.
[0241] Optionally, the communication module 1301 is further configured to send a persistence command to multiple non-faulty dedicated processors. The persistence command is used to instruct to persist the data sets included in the multiple non-faulty dedicated processors. For example, the communication module 1301 is configured to execute Figure 8 step 5120 in
[0242] Optionally, the integrity information includes the identification and correct information of the data in the data set; when the fault handling module 1305 determines that the data sets included in multiple non-faulty dedicated processors form updated optimizer data according to the integrity information sent by the non-faulty dedicated processors in the computing system, it is specifically used for: determining multiple consecutive identifications according to the identifications included in the integrity information sent by the non-faulty dedicated processors in the computing system, and the data indicated by the multiple consecutive identifications forms the updated optimizer data; determining that the data indicated by the multiple consecutive identifications is correct according to the correct information included in the integrity information; and determining the non-faulty dedicated processors where the data indicated by the multiple consecutive identifications is located as the multiple non-faulty dedicated processors that form the updated optimizer data.
[0243] Optionally, the integrity information further includes the number of iterations of the data in the data set included in the non-faulty dedicated processors; the fault handling module 1305 is further configured to determine that the number of iterations of the data indicated by the multiple consecutive identifications is the same according to the number of iterations included in the integrity information.
[0244] Optionally, the communication module 1301 is further configured to obtain the updated optimizer data when restarting model training; the optimizer data update module 1303 is configured to configure the updated first data set to the first dedicated processor.
[0245] Optionally, the communication module 1301 is further configured to configure the updated first data set to the third dedicated processor.
[0246] The storage module 1304 is used to store integrity information, data sets in optimizer data, etc.
[0247] It should be understood that the data processing device 1300 in the embodiments of the present application can be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. It can also be implemented by software Figure 5 、 Figure 8 、 Figure 9 When implementing the method shown, and its respective modules can also be software modules, and the computing graph optimization device 900 and its respective modules can also be software modules.
[0248] The data processing apparatus 1300 according to an embodiment of the present application may correspond to executing the methods described in the embodiments of the present application, and the above and other operations and / or functions of each unit in the data processing apparatus 1300 are respectively for implementing Figure 5 , Figure 8 , Figure 9 The corresponding processes of each method in, for the sake of brevity, will not be elaborated herein.
[0249] Figure 14 FIG. 10 is a schematic structural diagram of a computer device 1400 provided by the present application. As Figure 14 shown, the computer device 1400 includes a processor 1410, a bus 1420, a memory 1430, a communication interface 1440, a memory 1450 (which may also be referred to as a main memory unit), and a processor 1460. The processor 1410, the processor 1460, the memory 1430, the memory 1450, and the communication interface 1440 are connected through the bus 1420.
[0250] It should be understood that in this embodiment, the processor 1410 may be a CPU, and the processor 1410 may also be other general-purpose processors, digital signal processors (DSP), ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0251] The computer device 1400 may further include a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the solution of the present application. For example, the processor 1460 may be a GPU or an NPU.
[0252] The communication interface 1440 is used to implement the communication between the computer device 1400 and external devices or components.
[0253] In the present application, when the computer device 1400 is used to implement Figure 5 , Figure 8 , Figure 9 the functions of the dedicated processor shown, the communication interface 1440 is used to transmit gradients, initial values, etc., so that the processor 1460 can be used to train the model and update optimizer data, and persist the end-of-life checkpoint.
[0254] The computer device 1400 is used to implement Figure 5 , Figure 8 , Figure 9When referring to the functions of the general - purpose processor shown, the communication interface 1440 is used to transmit integrity information, etc., so that the processor 1410 can be used to indicate the persistence of data sets included in multiple non - faulty dedicated processors.
[0255] The bus 1420 may include a path for transmitting information between the above - mentioned components (such as the processor 1410, the memory 1450, and the storage 1430). In addition to the data bus, the bus 1420 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clear illustration, all kinds of buses are labeled as bus 1420 in the figure. The bus 1420 may be a Peripheral Component Interconnect Express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The bus 1420 can be divided into an address bus, a data bus, a control bus, etc.
[0256] As an example, the computer device 1400 may include multiple processors. The processor may be a multi - CPU processor. Here, the processor may refer to one or more devices, circuits, and / or computing units for processing data (such as computer program instructions).
[0257] It is worth noting that Figure 14 only the case where the computer device 1400 includes 1 processor 1410 and 1 storage 1430 is taken as an example. Here, the processor 1410 and the storage 1430 are respectively used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined according to service requirements. For example, the computer device 1400 includes multiple GPUs or NPUs.
[0258] The memory 1450 can be a volatile memory pool or a non-volatile memory pool, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). The memory 1450 is used to store optimizer data, gradients, and the like.
[0259] The memory 1430 can correspond to the storage medium for storing information such as optimizer data in the above method embodiments, for example, a disk, such as a mechanical hard disk or a solid-state drive.
[0260] The above computer device 1400 can be a general-purpose device or a special-purpose device. For example, the computer device 1400 can also be a server or other device with computing capabilities.
[0261] It should be understood that the computer device 1400 according to this embodiment can correspond to the data processing device 1300 in this embodiment, and can correspond to the corresponding entity executing any of the methods according to Figure 5 、 Figure 8 、 Figure 9 And the above and other operations and / or functions of each module in the data processing device 1300 are respectively for implementing Figure 5 、 Figure 8 、 Figure 9 The corresponding processes of each method in, for the sake of brevity, will not be described herein again.
[0262] The method steps in this embodiment can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), register, hard disk, removable hard disk, CD-ROM, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. Additionally, the ASIC can be located in a computing device. Of course, the processor and the storage medium can also exist as discrete components in the computing device.
[0263] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid state drive (SSD). The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A data processing method, characterized in that, applied to a computing system, the computing system includes a general - purpose processor and multiple dedicated processors, the multiple dedicated processors are used to train an artificial intelligence model, the multiple dedicated processors contain at least two copies of optimizer data, the optimizer data includes a first data set, the multiple dedicated processors include a first dedicated processor and a second dedicated processor, both the first dedicated processor and the second dedicated processor contain the first data set, and the method includes: The first dedicated processor trains the artificial intelligence model to obtain a first gradient; The first dedicated processor updates the first data set according to the first gradient to obtain an updated first data set; The second dedicated processor updates the first data set according to the first gradient to obtain the updated first data set; When the second dedicated processor fails, the general - purpose processor instructs to persist the data sets contained in multiple non - failed dedicated processors, the data sets contained in the multiple non - failed dedicated processors constitute the updated optimizer data, the multiple non - failed dedicated processors include the first dedicated processor, and the updated optimizer data includes the updated first data set.
2. The method according to claim 1, characterized in that, The general - purpose processor instructing to persist the data sets contained in multiple non - failed dedicated processors includes: Obtaining integrity information sent by non - failed dedicated processors in the computing system, the integrity information is used to indicate the integrity of the data sets contained in the non - failed dedicated processors; Determining that the data sets contained in the multiple non - failed dedicated processors constitute the updated optimizer data according to the integrity information sent by non - failed dedicated processors in the computing system; Sending a persistence command to the multiple non - failed dedicated processors, the persistence command is used to instruct to persist the data sets contained in the multiple non - failed dedicated processors.
3. The method according to claim 2, characterized in that, The integrity information includes the identifier and correct information of the data in the data set; Determining that the data sets contained in the multiple non - failed dedicated processors constitute the updated optimizer data according to the integrity information sent by non - failed dedicated processors in the computing system includes: Determining multiple consecutive identifiers according to the identifiers included in the integrity information sent by non - failed dedicated processors in the computing system, and the data indicated by the multiple consecutive identifiers constitutes the updated optimizer data; Determining that the data indicated by the multiple consecutive identifiers is correct according to the correct information included in the integrity information; Determining the non - failed dedicated processors where the data indicated by the multiple consecutive identifiers is located as the multiple non - failed dedicated processors that constitute the updated optimizer data.
4. The method according to claim 3, characterized in that, The integrity information further includes the iteration number of the data in the data sets contained in the non - failed dedicated processors; The method further includes: Determining that the iteration numbers of the data indicated by the multiple consecutive identifiers are the same according to the iteration number included in the integrity information.
5. The method according to any one of claims 1-4, wherein, the method further comprises: the first dedicated processor obtains an initial value of the first data set configured by the general-purpose processor; the second dedicated processor obtains the initial value of the first data set sent by the first dedicated processor.
6. The method according to any one of claims 1-5, wherein, the optimizer data further includes a second data set, the plurality of dedicated processors include a third dedicated processor and a fourth dedicated processor, and both the third dedicated processor and the fourth dedicated processor include the second data set.
7. The method according to any one of claims 1-6, wherein, after the general-purpose processor instructs to persist the data sets included in multiple non-faulty dedicated processors, the method further comprises: when restarting model training, the general-purpose processor obtains the updated optimizer data; the general-purpose processor configures the updated first data set to the first dedicated processor; the general-purpose processor configures the updated first data set to the third dedicated processor.
8. The method according to claim 7, wherein, the method further comprises: the first dedicated processor converts optimizer parameters in the updated first data set into model parameters; the first dedicated processor transmits the model parameters to a dedicated processor in the computing system for training the model parameters.
9. The method according to any one of claims 1-8, wherein, the method further comprises: the first dedicated processor transmits the first gradient to the second dedicated processor.
10. The method according to any one of claims 1-9, wherein, the optimizer data includes optimizer parameters and an optimizer state, and the optimizer state includes variance and momentum.
11. A computer device, wherein, the computer device includes a memory and a plurality of processors, the memory is used for storing a set of computer instructions; when the processors execute the set of computer instructions, the plurality of processors jointly execute the operation steps of the method according to any one of claims 1-10 above.
Citation Information
Cited By
Data processing method and computer device
EP4814723A1
Data processing method and computer device
WO2025119165A1