Data processing method and computer device

By deploying multi-copy optimizer data in the computing system and transmitting gradients between processors, the problem of processor failures resulting in incomplete optimizer data is solved, achieving high reliability of optimizer data and efficient model training.

WO2025119165A1PCT designated stage expired Publication Date: 2025-06-12HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/136412
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-12-03
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

When multi-processors collaboratively train artificial intelligence models, processor failures may lead to incomplete optimizer data, resulting in inaccurate checkpoint information, and wasting model training time.

Method used

Ensure the integrity of the optimizer data in the event of a processor failure by deploying at least two copies of optimizer data in a computing system and transmitting gradients between dedicated processors.

Benefits of technology

Improve the reliability of the optimizer data, ensure the accuracy of the end-of-life checkpoint, reduce the loss of model training, and improve training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136412_12062025_PF_FP_ABST
    Figure CN2024136412_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a data processing method and a computer device, which relate to the field of artificial intelligence. Globally, at least two dedicated processors include datasets in optimizer data, i.e., at least two copies of the optimizer data are deployed globally, thereby achieving a plurality of replicas of the optimizer data globally. Moreover, the datasets in the optimizer data included in the dedicated processors are updated by means of transmitting gradients between the dedicated processors, i.e., the plurality of replicas of the optimizer data are updated by means of "computation instead of transmission", thus ensuring the integrity of the optimizer data at all times. When a dedicated processor has a fault, since a system includes the plurality of copies of the optimizer data, complete optimizer data can be composed of the datasets included in non-faulty dedicated processors, thus ensuring the accuracy of a final checkpoint, so as to facilitate the resumption of model training on the basis of the final checkpoint and reduce training losses.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method and computer equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 7, 2023, with application number 202311690461.1 and application name “Data Processing Method and Computer Device,” all of the contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence, and in particular to a data processing method and computer equipment. Background Art

[0003] Currently, multiple processors are used to collaboratively train large parameter-rich artificial intelligence (AI) models. For example, large language models (LLMs) can be trained in a distributed manner across multiple machines and multiple graphics cards, where training can take up to several months. To prevent model training interruptions due to processor, network, or software failures, checkpoint information can be periodically saved during model training. For example, checkpoint information includes model parameters and optimizer data.

[0004] However, each processor stores a portion of the optimizer data, and each processor stores a different portion of the data. There is only one global copy, and the data stored by multiple processors can form the complete optimizer data. If a processor fails, some of the optimizer data stored by that processor may be lost, resulting in incomplete optimizer data, inaccurate checkpoint information, and wasted model training time. Summary of the Invention

[0005] The present application provides a data processing method and computer device, thereby ensuring the integrity of optimizer data when a processor fails.

[0006] In a first aspect, a data processing method is provided. The computing system used in the method includes a general-purpose processor and multiple special-purpose processors, wherein the multiple special-purpose processors are used to train artificial intelligence models. The multiple special-purpose processors contain at least two copies of optimizer data, the optimizer data including a first data set, and at least two special-purpose processors each contain the first data set. For example, the multiple special-purpose processors include a first special-purpose processor and a second special-purpose processor, each of which contains the first data set. The method includes: the first special-purpose processor trains the artificial intelligence model to obtain a first gradient; the first special-purpose processor updates the first data set based on the first gradient to obtain an updated first data set; and the second special-purpose processor updates the first data set based on the first gradient to obtain an updated first data set. When the second special-purpose processor fails, the general-purpose processor instructs the persistence of the data sets contained in the multiple non-faulty special-purpose processors, the data sets contained in the multiple non-faulty special-purpose processors forming the updated optimizer data, the multiple non-faulty special-purpose processors including the first special-purpose processor, and the updated optimizer data including the updated first data set.

[0007] Compared with the global system containing one copy of optimizer data, a failure of a dedicated processor results in an inaccurate terminal checkpoint and incomplete optimizer data. The method provided by the present application is that at least two dedicated processors in the global system contain a data set in the optimizer data, that is, at least two copies of optimizer data are deployed in the global system to achieve multiple copies of optimizer data in the global system. In addition, the dedicated processors update the data set in the optimizer data contained in the dedicated processor by transmitting gradients, that is, the update of multiple copies of optimizer data is achieved by "using calculation instead of transmission", thereby ensuring the integrity of the optimizer data at any time. When a dedicated processor fails, since the system contains multiple copies of optimizer data, the complete optimizer data can be formed by the data set contained in the non-faulty dedicated processors, thereby improving the reliability of the optimizer data and ensuring the accuracy of the terminal checkpoint, so as to facilitate the recovery of model training based on the terminal checkpoint and reduce training losses.

[0008] In a possible implementation, the method further includes: the first dedicated processor acquiring an initial value of a first data set configured by the general processor; and the second dedicated processor acquiring the initial value of the first data set sent by the first dedicated processor.

[0009] Thereby, it is ensured that at least two dedicated processors contain the same data set.

[0010] In another possible implementation, the method further includes: the first dedicated processor transmitting the first gradient to the second dedicated processor.

[0011] This allows dedicated processors with the same dataset to retrieve gradients and update the dataset, ensuring that at least two dedicated processors have the same updated dataset. Because optimizer data is large and gradients are small, transmitting gradients can reduce bandwidth and the amount of data transmitted.

[0012] In another possible implementation, the optimizer data further includes a second data set, the plurality of dedicated processors include a third dedicated processor and a fourth dedicated processor, and the third dedicated processor and the fourth dedicated processor both include the second data set.

[0013] It is understandable that the first data set contained in the first dedicated processor and the second data set contained in the third dedicated processor constitute the complete optimizer data. The first data set contained in the second dedicated processor and the second data set contained in the fourth dedicated processor constitute the complete optimizer data.

[0014] Therefore, at least two dedicated processors in the global system contain data sets in the optimizer data, that is, at least two copies of the optimizer data are deployed in the global system, thus realizing multiple copies of the optimizer data in the global system.

[0015] In another possible implementation, the method also includes: a third dedicated processor training the artificial intelligence model to obtain a second gradient; the third dedicated processor updating the second data set according to the second gradient to obtain an updated second data set; a fourth dedicated processor obtaining the second gradient from the third dedicated processor; the fourth dedicated processor updating the second data set according to the second gradient to obtain an updated second data set.

[0016] Therefore, by "using calculation instead of transmission", it is possible to update multiple copies of optimizer data and ensure the integrity of optimizer data at all times.

[0017] In another possible implementation, a general-purpose processor instructs persistence of a data set included in multiple non-faulty dedicated processors, including: obtaining integrity information sent by non-faulty dedicated processors in a computing system, the integrity information being used to indicate the integrity of the data set included in the non-faulty dedicated processors; determining, based on the integrity information sent by the non-faulty dedicated processors in the computing system, that the data sets included in the multiple non-faulty dedicated processors constitute updated optimizer data; and sending a persistence command to the multiple non-faulty dedicated processors. The persistence command is used to instruct persistence of the data set included in the multiple non-faulty dedicated processors.

[0018] In this way, the general-purpose processor determines the dedicated processor that can form the complete optimizer data based on the integrity information of the data contained in the dedicated processor, ensuring the integrity of the optimizer data when the dedicated processor fails, improving the reliability of the optimizer data, and ensuring the accuracy of the dying checkpoint, so as to facilitate the resumption of model training based on the dying checkpoint and reduce training losses.

[0019] In another possible implementation, the integrity information includes the identification and correctness information of the data in the data set; determining that the data sets contained in multiple non-faulty dedicated processors constitute the updated optimizer data based on the integrity information sent by the non-faulty dedicated processors in the computing system include: determining multiple consecutive identifiers based on the identifiers included in the integrity information sent by the non-faulty dedicated processors in the computing system, the data indicated by the multiple consecutive identifiers constitute the updated optimizer data; determining that the data indicated by the multiple consecutive identifiers are correct based on the correctness information included in the integrity information; and determining the non-faulty dedicated processors where the data indicated by the multiple consecutive identifiers are located as the multiple non-faulty dedicated processors constituting the updated optimizer data.

[0020] In another possible implementation, the integrity information further includes the number of iterations of data in the data set included in the non-faulty dedicated processor; the method further includes: determining, based on the number of iterations included in the integrity information, that the number of iterations of data indicated by multiple consecutive identifiers is the same.

[0021] In this way, the number of data iterations ensures that the complete optimizer data is composed based on the latest data.

[0022] In another possible implementation, after the general-purpose processor instructs to persist the data sets contained in multiple non-faulty dedicated processors, the method also includes: when restarting model training, the general-purpose processor obtains updated optimizer data; the general-purpose processor configures the updated first data set to the first dedicated processor; and the general-purpose processor configures the updated first data set to the third dedicated processor.

[0023] In another possible implementation, the method further includes: the first dedicated processor converting the updated first data center optimizer parameters into model parameters; and the first dedicated processor transmitting the model parameters to a dedicated processor in the computing system for training the model parameters.

[0024] Since the dying checkpoint only saves the optimizer data and not the model parameters, when resuming model training, the model parameters are calculated based on the optimizer data in the storage system and then transmitted to the dedicated processor used to train the model parameters, thereby ensuring that the optimizer data and model parameters at the dying checkpoint are consistent. This overcomes the problem of inconsistency between the optimizer data and model parameters due to the failure of the AllGather operation when the dedicated processor fails. In addition, the method provided in this application can save the checkpoint at the time of the failure and resume model training based on the checkpoint at the time of the failure, thereby effectively reducing training losses and improving the efficiency of model training.

[0025] In another possible implementation, the optimizer data includes optimizer parameters and optimizer state, and the optimizer state includes variance and momentum.

[0026] In a second aspect, a data processing method is provided. The computing system used in the method includes a general-purpose processor and multiple special-purpose processors, and the multiple special-purpose processors are used to train artificial intelligence models. The multiple special-purpose processors contain at least two optimizer data, the optimizer data includes a first data set, and at least two special-purpose processors each contain the first data set. For example, the multiple special-purpose processors include a first special-purpose processor and a second special-purpose processor, and the first special-purpose processor and the second special-purpose processor each contain the first data set. The method is executed by the first special-purpose processor and includes: training the artificial intelligence model to obtain a first gradient; updating the first data set based on the first gradient to obtain an updated first data set; and transmitting the first gradient to the second special-purpose processor.

[0027] In a possible implementation manner, the method further includes: obtaining an initial value of a first data set configured by the general processor.

[0028] In another possible implementation, the method further includes: converting the updated optimizer parameters in the first data set into model parameters; and transmitting the model parameters to a dedicated processor in the computing system for training the model parameters.

[0029] In another possible implementation, the method further includes: persisting the included data set.

[0030] In a third aspect, a data processing method is provided. The computing system used in the method includes a general-purpose processor and multiple special-purpose processors, and the multiple special-purpose processors are used to train artificial intelligence models. The multiple special-purpose processors contain at least two optimizer data, the optimizer data includes a first data set, and at least two special-purpose processors each contain the first data set. For example, the multiple special-purpose processors include a first special-purpose processor and a second special-purpose processor, and the first special-purpose processor and the second special-purpose processor each contain the first data set. The method is executed by the second special-purpose processor and includes: obtaining a first gradient sent by the first special-purpose processor; and updating the first data set based on the first gradient to obtain an updated first data set.

[0031] In a possible implementation manner, the method further includes: acquiring an initial value of the first data set sent by the first dedicated processor.

[0032] In another possible implementation, the method further includes: converting the updated optimizer parameters in the first data set into model parameters; and transmitting the model parameters to a dedicated processor in the computing system for training the model parameters.

[0033] In another possible implementation, the method further includes: persisting the included data set.

[0034] In another possible implementation, the method further includes: training an artificial intelligence model to obtain a second gradient.

[0035] In a fourth aspect, a data processing method is provided, wherein the computing system used by the method includes a general-purpose processor and multiple special-purpose processors, and the multiple special-purpose processors are used to train artificial intelligence models. The multiple special-purpose processors contain at least two copies of optimizer data, the optimizer data includes a first data set, and at least two special-purpose processors each contain the first data set. For example, the multiple special-purpose processors include a first special-purpose processor and a second special-purpose processor, and the first special-purpose processor and the second special-purpose processor each contain a first data set. The method is executed by a general-purpose processor, and the method includes: when the second special-purpose processor fails, instructing to persist the data sets contained in the multiple non-faulty special-purpose processors, the data sets contained in the multiple non-faulty special-purpose processors constitute the updated optimizer data, the multiple non-faulty special-purpose processors include the first special-purpose processor, and the updated optimizer data includes the updated first data set.

[0036] In one possible implementation, a general-purpose processor instructs persistence of a data set contained in multiple non-faulty dedicated processors, including: obtaining integrity information sent by non-faulty dedicated processors in a computing system, the integrity information being used to indicate the integrity of the data set contained in the non-faulty dedicated processors; determining that the data sets contained in multiple non-faulty dedicated processors constitute updated optimizer data based on the integrity information sent by non-faulty dedicated processors in the computing system; and sending a persistence command to the multiple non-faulty dedicated processors, the persistence command being used to instruct persistence of the data sets contained in the multiple non-faulty dedicated processors.

[0037] In another possible implementation, the integrity information includes the identification and correctness information of the data in the data set; determining that the data sets contained in multiple non-faulty dedicated processors constitute the updated optimizer data based on the integrity information sent by the non-faulty dedicated processors in the computing system includes: determining multiple consecutive identifiers based on the identifiers contained in the integrity information sent by the non-faulty dedicated processors in the computing system, the data indicated by the multiple consecutive identifiers constitute the updated optimizer data; determining that the data indicated by the multiple consecutive identifiers are correct based on the correctness information contained in the integrity information; and determining the non-faulty dedicated processors where the data indicated by the multiple consecutive identifiers are located as the multiple non-faulty dedicated processors constituting the updated optimizer data.

[0038] In another possible implementation, the integrity information further includes the number of iterations of data in the data set included in the non-faulty dedicated processor; the method further includes: determining, based on the number of iterations included in the integrity information, that the number of iterations of data indicated by multiple consecutive identifiers is the same.

[0039] In another possible implementation, the method further includes: when restarting model training, obtaining updated optimizer data; configuring the updated first data set to the first dedicated processor; and configuring the updated first data set to the third dedicated processor.

[0040] In a fifth aspect, a data processing device is provided, the data processing device including modules of a dedicated processor for executing the method of the first aspect or any possible design of the first aspect. For example, the data processing device includes a communication module, a training module, and an optimizer data update module.

[0041] The training module is configured to train the artificial intelligence model to obtain a first gradient. The optimizer data update module is configured to update the first data set according to the first gradient to obtain an updated first data set.

[0042] In a possible implementation, the communication module is configured to transmit the first gradient.

[0043] In another possible implementation, the communication module is further configured to obtain an initial value of a first data set configured for the general processor.

[0044] In another possible implementation, the optimizer data update module is further configured to convert the updated optimizer parameters in the first data set into model parameters. The communication module is further configured to transmit the model parameters to a dedicated processor in the computing system for training the model parameters.

[0045] In another possible implementation, the optimizer data update module is also used to persist the included data sets.

[0046] In a sixth aspect, a data processing device is provided, the data processing device including modules of a dedicated processor for executing the method of the first aspect or any possible design of the first aspect. For example, the data processing device includes a communication module, a training module, and an optimizer data update module.

[0047] The communication module is used to obtain a first gradient; the optimizer data update module is used to update the first data set according to the first gradient to obtain an updated first data set.

[0048] In a possible implementation, the communication module is further configured to obtain an initial value of the first data set.

[0049] In another possible implementation, the optimizer data update module is further configured to convert the updated optimizer parameters in the first data set into model parameters. The communication module is further configured to transmit the model parameters to a dedicated processor in the computing system for training the model parameters.

[0050] In another possible implementation, the optimizer data update module is also used to persist the included data sets.

[0051] In a seventh aspect, a data processing device is provided, the data processing device including modules of a general-purpose processor for executing the method of the first aspect or any possible design of the first aspect. For example, the data processing device includes a communication module, a fault handling module, and an optimizer data update module.

[0052] A fault handling module is used to instruct the persistence of data sets contained in multiple non-faulty dedicated processors when a dedicated processor in the system fails. The data sets contained in the multiple non-faulty dedicated processors constitute updated optimizer data. The multiple non-faulty dedicated processors include a first dedicated processor, and the updated optimizer data includes the updated first data set.

[0053] In one possible implementation, the communication module is configured to obtain integrity information sent by a non-faulty dedicated processor in the computing system, where the integrity information is configured to indicate integrity of a data set contained in the non-faulty dedicated processor;

[0054] When the fault handling module instructs to persist the data sets contained in multiple non-faulty dedicated processors, it is specifically used to: determine the updated optimizer data composed of the data sets contained in multiple non-faulty dedicated processors based on the integrity information sent by the non-faulty dedicated processors in the computing system; the communication module is also used to send a persistence command to the multiple non-faulty dedicated processors, and the persistence command is used to instruct to persist the data sets contained in multiple non-faulty dedicated processors.

[0055] In another possible implementation, the integrity information includes the identification and correctness information of the data in the data set; when the fault processing module determines that the data sets contained in multiple non-faulty dedicated processors constitute the updated optimizer data based on the integrity information sent by the non-faulty dedicated processors in the computing system, it is specifically used to: determine multiple consecutive identifiers based on the identifiers contained in the integrity information sent by the non-faulty dedicated processors in the computing system, and the data indicated by the multiple consecutive identifiers constitute the updated optimizer data; determine that the data indicated by the multiple consecutive identifiers are correct based on the correctness information contained in the integrity information; and determine the non-faulty dedicated processors where the data indicated by the multiple consecutive identifiers are located as the multiple non-faulty dedicated processors constituting the updated optimizer data.

[0056] In another possible implementation, the integrity information further includes the number of iterations of data in the data set included in the non-faulty dedicated processor; the fault processing module is further used to determine, based on the number of iterations included in the integrity information, that the number of iterations of data indicated by multiple consecutive identifiers is the same.

[0057] In another possible implementation, the communication module is also used to obtain updated optimizer data when restarting model training; the optimizer data update module is used to configure the updated first data set to the first dedicated processor; the communication module is also used to configure the updated first data set to the third dedicated processor.

[0058] In an eighth aspect, a computing system is provided, which includes a general-purpose processor and multiple special-purpose processors, and a memory is used to store a set of computer instructions; when the processor executes the set of computer instructions, the general-purpose processor and the multiple special-purpose processors jointly execute the operating steps of the method in the first aspect or any possible implementation of the first aspect.

[0059] In the ninth aspect, a computer device is provided, which includes a memory and multiple processors, the memory being used to store a set of computer instructions; when the processor executes the set of computer instructions, the multiple processors jointly execute the operating steps of the method in the first aspect or any possible implementation of the first aspect.

[0060] In the tenth aspect, a computer-readable storage medium is provided, comprising: computer software instructions; when the computer software instructions are executed in a processor, the processor executes the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0061] In the eleventh aspect, a computer program product is provided. When the computer program product is run on a computer, it enables the computer to perform the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0062] The technical effects brought about by any design method in the second aspect to the eleventh aspect can be referred to the technical effects brought about by the first aspect or different design methods in the first aspect, and will not be repeated here.

[0063] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is a schematic diagram of a large model training provided by this application;

[0065] FIG2 is a schematic diagram of a checkpoint provided by the present application;

[0066] FIG3 is a schematic diagram of a logical process of model training provided by this application;

[0067] FIG4 is a schematic diagram of the architecture of a data processing system provided by the present application;

[0068] FIG5 is a flow chart of a data processing method provided by the present application;

[0069] FIG6 is a flow chart of a data processing method provided by the present application;

[0070] FIG7 is a schematic diagram of a gradient transmission provided by the present application;

[0071] FIG8 is a schematic diagram of a dying checkpoint preservation method provided by the present application;

[0072] FIG9 is a flow chart of a data processing method provided by the present application;

[0073] FIG10 is a schematic diagram of an optimizer data storage provided by the present application;

[0074] FIG11 is a schematic structural diagram of a data processing device provided by the present application;

[0075] FIG12 is a schematic structural diagram of a data processing device provided by the present application;

[0076] FIG13 is a schematic structural diagram of a data processing device provided by the present application;

[0077] FIG14 is a schematic structural diagram of a computer device provided in this application. DETAILED DESCRIPTION

[0078] To facilitate understanding, the main terms involved in this application are first explained.

[0079] Large models: refers to ultra-large-scale artificial intelligence (AI) models. Large models are widely used in the field of natural language processing and are completely changing the status of natural language processing (NLP) tasks, giving rise to more powerful and intelligent language technologies. Large models are one of the important directions of AI development. Large models also have the ability to perform well in various natural language processing tasks, such as text classification, sentiment analysis, summary generation, translation, etc. Large models can be used in multiple application areas such as automatic writing, chatbots, virtual assistants, voice assistants, automatic translation, etc. For example, large language models (LLM), bidirectional encoder representations from transformers (BERT), generative pre-trained transformer models (GPT), GPT3, GPT4, MoE, etc. Large models have the following characteristics.

[0080] 1. Huge scale: Model size can reach hundreds of gigabytes (GB) or even larger. Large models contain billions, hundreds of billions, or even trillions of model parameters. This huge scale provides powerful expressive and learning capabilities.

[0081] 2. Multi-task learning: Large models can handle a variety of different natural language processing (NLP) tasks, such as machine translation, text summarization, and question-answering systems, enabling the model to learn broader and more generalized language understanding capabilities.

[0082] 3. Powerful computing resources: Training large models typically requires hundreds or even thousands of specialized processors and a significant amount of time. For example, training a large model can take weeks to months. Powerful computing resources can accelerate the training process while preserving the performance of large models. Dedicated processors include but are not limited to graphics processing units (GPUs), data processing units (DPUs), neural processing units (NPUs), and embedded neural-network processing units (NPUs).

[0083] 4. Rich data: Use a large amount of training data to train large models and give full play to the scale advantages of the model parameters of large models.

[0084] Model parameters: These are the variables or weights that need to be learned or adjusted in an AI model. Model parameters can affect the model's predictive power and performance. During model training, the model optimizes its performance by trying different parameter combinations. Common model parameters include weights, biases, learning rates, and regularization coefficients. When the model is used for prediction, these model parameters are used to calculate the output.

[0085] Model training involves using a training set to train an AI model, enabling it to predict or classify unknown data. During model training, the model learns based on the features and target values ​​in the training set. Upon completion, a model is generated that can be used to predict or classify unknown data. Model training is one of the most critical steps in machine learning, directly impacting the model's accuracy and reliability.

[0086] Optimizer: It can refer to an algorithm used to adjust the model parameters of an artificial intelligence model so that the artificial intelligence model can predict the output results more accurately. The purpose of the optimizer is to minimize the loss function, that is, the difference between the model's predicted value and the actual value. Common optimizers include stochastic gradient descent (SGD), adaptive moment estimation (Adam), adaptive gradient (Adagrad), and root mean square prop (RMSprop). These optimizers use different strategies to update model parameters to enable the artificial intelligence model to achieve better performance and accuracy. Usually the input of the optimizer is the gradient and the output is the model parameters.

[0087] Checkpoint (CKPT): refers to saving checkpoint data at regular intervals or training rounds during model training so that the model can be restored and continued to be trained later. Checkpoints include model parameters and optimizer data. Optimizer data includes optimizer parameters and optimizer state (OS). Optimizer state includes variance and momentum. For example, if an unexpected situation occurs during model training and the model training is interrupted, the model training can be restored based on the most recent checkpoint without having to restart the model training. Checkpoints can also be used for model evaluation and debugging. For example, the performance of the model at different training stages can be observed based on different checkpoints to determine the model training effect and optimization direction. In deep learning, checkpoints are a very important way to save and manage model parameters.

[0088] Resumable training is a deep learning training technique that allows you to pause model training and save the current data. This data typically includes model parameters and optimizer data. When you restart training, the model can continue from where it was last paused, without having to start from the beginning. This technique can accelerate model training, reduce the waste of computing resources, and handle various issues that may arise during model training, such as computer failures or network outages.

[0089] Figure 1 is a schematic diagram of a large model training method provided by this application. As shown in Figure 1, computing system 100 includes multiple cards. Cards can also be called training cards, accelerator cards, or dedicated processors. Multiple cards can be located in multiple servers. The servers can be called hosts. Parallel technology is used based on multiple cards to accelerate the training of large models. Parallel technologies include data parallelism, model parallelism, pipeline parallelism, and optimizer parallelism.

[0090] Data parallelism involves dividing multiple cards into multiple data parallel domains. As shown in Figure 1, multiple cards are divided into X data parallel domains. The training set is split based on the data parallel domains, and the cards contained in each data parallel domain train a large model based on a subset of the training set. Multiple data parallel domains can use different subsets of the training set to train the large model in parallel.

[0091] Model parallelism involves splitting a large model into multiple sub-models. For example, a large model can be split into multiple layers, with each card capable of training at least one layer. Alternatively, if a single layer of a model is large, this layer can be split into multiple cards capable of training that layer. Different data parallel domains can train different layers of a large model. Multiple data parallel domains can train a large model in parallel.

[0092] Pipeline parallelism refers to dividing a large model into multiple layers based on the logical order of the layers within the model. Multiple cards can then train the layers of the large model in parallel or serially, based on the logical order of the layers within the large model. The logical order between layers in a large model can refer to the dependencies between them. For example, if the output data of the first layer is the input data of the second layer, two cards can train the first and second layers serially.

[0093] Optimizer parallelism divides optimizer data into multiple datasets based on the number of cards in the computing system. Each card contains a portion of the optimizer data, reducing the optimizer data storage requirements on the cards. Each card stores a different dataset, and there is only one global copy of a dataset. Datasets stored by multiple cards can form the complete optimizer data.

[0094] For example, during LLM training, the data stored in the card mainly includes activation, model parameters, and optimizer data. Activation and model parameters are represented by half-precision floating-point (FP16), and optimizer data is represented by single-precision floating point (e.g., FP32). Assuming that the number of model parameters is M, the storage space required for the model parameters is 2M Bytes, and the storage space required for the optimizer data is 12M Bytes. That is, the optimizer data includes M optimizer parameters, M variances, and M momenta. The storage space required for M optimizer parameters is 4M Bytes. The storage space required for M variances is 4M Bytes. The storage space required for M momenta is 4M Bytes. For example, GPT3 includes 175 billion model parameters, and the storage space required for optimizer data is 12 bytes*175 billion=2.45 terabytes (TB). In the distributed training LLM, multiple machines and multiple cards work together. Each card stores a portion of the optimizer data, and multiple cards store the complete optimizer data.

[0095] Among them, each card performs forward, activation, reverse, gradient generation, gradient accumulation and other operations, and distributes the gradient to multiple cards through the AllReduce operation. The card updates the optimizer data with the same identifier as the gradient according to the gradient; the model parameters are obtained according to the optimizer parameter conversion through the AllGather operation, thereby realizing model training.

[0096] Furthermore, the global model contains multiple copies of model parameters. Each data-parallel domain contains a complete copy of the model parameters. Cards in different data-parallel domains contain the same model parameters. For example, data-parallel domain 1 contains M model parameters, and data-parallel domain 2 contains M model parameters. These M model parameters are distributed and stored across the cards in the data-parallel domains. For example, card 1, card 5, and card x contain parameter 1. Card 2, card 6, and card x+1 contain parameter 2.

[0097] The global system contains one copy of the optimizer data. Assume the system has n cards, and each card contains 1 / n of the optimizer data. The n copies of 1 / n of the optimizer data contained on n cards constitute the complete optimizer data.

[0098] In some embodiments, the card can save checkpoints periodically, i.e., save a checkpoint every once in a while. This checkpoint can be called a periodic checkpoint. If model training fails at any point in time, model training can be resumed based on the previous checkpoint.

[0099] For example, as shown in Figure 2 (a), card n fails. Cards 1 and 2 persist model parameters and optimizer data to the storage system. As shown in Figure 2 (b), during model training, the i-th checkpoint and the i+1-th checkpoint are saved. After the i+1-th checkpoint, card n fails, interrupting model training. The failure is detected, the failure mode is determined, and whether training should be restarted is determined. If model training is restarted, it is resumed based on the i+1-th checkpoint.

[0100] Since the model training is resumed from the last checkpoint, some training is lost. For example, the optimizer data and model parameter update data during the training period from the i+1th checkpoint to the failure point are lost.

[0101] Measuring training loss primarily involves two factors: 1. The number of GPUs used for training. Training large models requires thousands of GPUs, and even minute-level losses combined with the number of GPUs multiplied by the loss time can result in significant training loss. 2. The probability of GPU and network failures. A high GPU or network failure rate increases the frequency of training interruptions, leading to greater training loss.

[0102] The training loss can be used to calculate the effectiveness of multiple cards. The higher the training loss, the lower the effectiveness of the card.

[0103] There are two main ways to reduce training loss: 1. Reduce the probability of failure; 2. When a failure occurs, try to recover from the last state. For example, try saving a checkpoint at the time of the failure and resuming model training based on the checkpoint at the time of the failure to reduce training loss. The checkpoint at the time of the failure can be called a final checkpoint.

[0104] There are two issues regarding how to implement the technology for the end-of-life checkpoint:

[0105] 1. Checkpoint integrity: As shown in Figure 1, the global system contains only one copy of optimizer data, and each card contains 1 / n of the optimizer data. Once a card fails, the checkpoint is incomplete, which also leads to incomplete optimizer data.

[0106] 2. Checkpoint data consistency: Due to incomplete optimizer data, model parameters cannot be updated based on the optimizer data. This leads to inconsistencies between the optimizer data and model parameters, and slows down the convergence of model training. In fact, it is even worse than resuming model training based on the last periodic checkpoint.

[0107] The situations in which the optimizer data and model parameters are inconsistent mainly include that the optimizer data is the latest optimizer data, the model parameters are the model parameters of the previous iteration, or the model parameters on some cards are the latest model parameters, but the model parameters on other cards are the model parameters of the previous iteration. The main reason for this situation is that the AllGather operation failed. For example, as shown in Figure 3, a logical process diagram of model training provided by this application is provided. The gradient is obtained through the AllReduce operation, the optimizer data is updated according to the gradient, and the optimizer data of checkpoint 1 is saved. The optimizer parameters are converted into model parameters, for example, the optimizer parameter fp32 is converted into the model parameter fp16, but due to a card failure, the AllGather operation fails and the model parameters of checkpoint 2 cannot be updated, resulting in inconsistency between the optimizer data and the model parameters. Among them, the data volume of the optimizer data is 12M. Taking M=175B as an example, the size of the optimizer data is 2.4TB and the size of the model parameters is 400GB.

[0108] To address the issue of inaccurate end-of-life checkpoints and incomplete optimizer data caused by a dedicated processor failure, the present application provides a data processing method. The computing system used in the method includes a general-purpose processor and multiple dedicated processors, and the multiple dedicated processors are used to train artificial intelligence models. The multiple dedicated processors contain at least two copies of optimizer data, and the optimizer data includes a first data set. At least two dedicated processors each contain the first data set. For example, the multiple dedicated processors include a first dedicated processor and a second dedicated processor, and the first and second dedicated processors each contain the first data set. The method includes: the first dedicated processor trains the artificial intelligence model to obtain a first gradient; the first dedicated processor updates the first data set based on the first gradient to obtain an updated first data set; the second dedicated processor updates the first data set based on the first gradient to obtain an updated first data set. When the second dedicated processor fails, the general-purpose processor instructs the persistence of the data sets contained in the multiple non-faulty dedicated processors. The data sets contained in the multiple non-faulty dedicated processors constitute the updated optimizer data. The multiple non-faulty dedicated processors include the first dedicated processor, and the updated optimizer data includes the updated first data set.

[0109] Compared with the global system containing one copy of optimizer data, a failure of a dedicated processor results in an inaccurate terminal checkpoint and incomplete optimizer data. The method provided by the present application is that at least two dedicated processors in the global system contain a data set in the optimizer data, that is, at least two copies of optimizer data are deployed in the global system to achieve multiple copies of optimizer data in the global system. In addition, the dedicated processors update the data set in the optimizer data contained in the dedicated processor by transmitting gradients, that is, the update of multiple copies of optimizer data is achieved by "using calculation instead of transmission", thereby ensuring the integrity of the optimizer data at any time. When a dedicated processor fails, since the system contains multiple copies of optimizer data, the complete optimizer data can be formed by the data set contained in the non-faulty dedicated processors, thereby improving the reliability of the optimizer data and ensuring the accuracy of the terminal checkpoint, so as to facilitate the recovery of model training based on the terminal checkpoint and reduce training losses.

[0110] The implementation of the data processing method provided in this application is described in detail below with reference to the accompanying drawings.

[0111] FIG4 is a schematic diagram of the architecture of a data processing system provided by the present application. As shown in FIG4 , a data processing system 400 includes a client 410 , a computing cluster 420 , and a storage cluster 430 .

[0112] Storage cluster 430 includes multiple storage nodes 431. A storage node 431 includes one or more controllers, a network card, and multiple hard disks. Hard disks are used to store data. Hard disks can be magnetic disks or other types of storage media, such as solid-state drives or shingled magnetic recording hard disks. The network card is used to communicate with computing nodes 421 included in computing cluster 420. The controller is used to write data to or read data from the hard disks based on read / write data requests sent by computing nodes 421. During the data reading and writing process, the controller needs to convert the addresses carried in the read / write data requests into addresses that the hard disks can recognize.

[0113] The computing cluster 420 includes multiple computing nodes 421. The computing node 421 can be a computing device, such as an accelerator card, a server, etc.

[0114] In some embodiments, computing cluster 420 may be a heterogeneous computing architecture to provide high-performance computing. For example, computing node 421 may include computing units with computing capabilities, such as a central processing unit (CPU), a graphics processing unit (GPU), a data processing unit (DPU), a neural processing unit (NPU), and an embedded neural-network processing unit (NPU), to provide high-performance computing.

[0115] In other embodiments, multiple computing nodes 421 are connected through network devices (such as switches, network cards, etc.) based on high-speed interconnection technology, so that the multiple computing nodes 421 can communicate with each other.

[0116] Client 410 communicates with computing cluster 420 and storage cluster 430 via network 440. For example, client 410 sends a request to computing cluster 420 via network 440, requesting that computing cluster 420 perform model training. Network 440 can be an internal enterprise network (e.g., a local area network (LAN)) or the Internet. Client 410 refers to a computer connected to network 440, also known as a workstation. Different clients can share network resources (e.g., computing resources and storage resources).

[0117] In some embodiments, computing cluster 420 also includes a control node 422. For example, the control node and the computing nodes can be independent physical devices. In another example, the control node and multiple computing nodes can be located on the same physical device. The control node can be a CPU. The multiple computing nodes include computing units such as GPUs, NPUs, and DPUs. Control node 422 is used to manage and allocate tasks, allowing multiple computing nodes to execute multiple tasks in parallel to increase data processing speed.

[0118] In the present application, the control node 422 is used to instruct multiple computing nodes to perform model training based on the parallel technology described in the above embodiment according to the request.

[0119] The control node 422 is also used to divide multiple computing nodes in the system into multiple groups, each group contains at least two computing nodes, and each group contains a complete optimizer data to achieve global multiple copies of optimizer data.

[0120] The control node 422 is further configured to instruct a non-faulty computing node to persist the optimizer data when a computing node fails.

[0121] Gradients may be transmitted between different groups of computing nodes 421 so that the computing nodes 421 update the data sets in the optimizer data according to the gradients.

[0122] In the embodiment of the present application, the storage cluster 430 stores optimizer data and the like.

[0123] In other embodiments, client 410 is installed with client program 411. Client 410 runs client program 411 to display a user interface (UI). User 450 operates the UI to submit a request. For example, user 450 operates the UI to submit a request. After receiving the request, control node 422 loads optimizer data from storage cluster 430 and distributes the optimizer data to multiple computing nodes, enabling the multiple computing nodes to perform model training based on the parallel technology described in the above embodiments.

[0124] Optionally, the system administrator 460 can call the application platform interface (API) 412 or the command-line interface (CLI) interface 413 through the client 410 to configure system information, such as the initial value of the optimizer data configured for the computing node provided in this application.

[0125] FIG4 is merely a schematic diagram. The embodiments of this application do not limit the device connection method or number of devices in the data processing system. For example, the data processing system may include multiple clients. A client may connect to multiple computing nodes. Different clients may establish connections with different computing nodes.

[0126] Next, the data processing process is described in detail with reference to the accompanying drawings.

[0127] FIG5 is a flow chart of a data processing method provided by this application. Here, the optimizer data update and the end-of-life checkpoint are mainly described. Assume that the computing system includes a general-purpose processor and K dedicated processors.

[0128] As shown in (a) of FIG5 , the dedicated processor performs an initialization process before training the model. The method includes the following steps 510 to 530 .

[0129] Step 510: The general processor obtains a request.

[0130] The general-purpose processor receives a request from the client, which is used to instruct the execution of model training. The request may include a model identifier, so that the general-purpose processor can identify the model based on the model identifier and instruct the dedicated processor to load the model indicated by the model identifier from the storage system.

[0131] The general-purpose processor can obtain optimizer data from the storage system. It should be noted that when the general-purpose processor instructs the dedicated processor to train the model for the first time, the general-purpose processor can obtain the initial optimizer data from the storage system. The initial optimizer data includes the initial values ​​of the optimizer parameters, the initial values ​​of the momentum, and the initial values ​​of the variance. When the general-purpose processor instructs the dedicated processor to resume model training, that is, when the dedicated processor performs breakpoint resume training, the general-purpose processor can obtain the updated optimizer data from the storage system. The updated optimizer data includes the updated values ​​of the optimizer parameters, the updated values ​​of the momentum, and the updated values ​​of the variance.

[0132] Step 520: The general processor initializes optimizer data.

[0133] The general-purpose processor can group the dedicated processors in the system into multiple groups. Each group contains at least two dedicated processors. The number of dedicated processors in each group can be the same or different.

[0134] General-purpose processors divide the optimizer data based on the number of dedicated processors in the group. Each dedicated processor in the group contains a portion of the optimizer data. At least two dedicated processors in different groups contain the same optimizer data. Each group contains a complete copy of the optimizer data, and multiple groups contain multiple copies of the complete optimizer data, achieving global multiple copies of the optimizer data.

[0135] For example, a general-purpose processor divides K specialized processors in the system into R groups. Each group contains at least two specialized processors. The value of R is an integer greater than or equal to 2. The larger the number of groups R, the more replicas of optimizer data there are in the system, making it easier to obtain complete optimizer data, thereby improving optimizer data reliability.

[0136] For ease of description, the following takes an example where each group includes N dedicated processors.

[0137] The general-purpose processor selects one of the R groups and divides the optimizer data into N shares based on the number of specialized processors N in the group. Each of the N specialized processors in the group contains N shares of data. That is, each specialized processor in the group contains 1 / N of the optimizer data. The N shares of data contained by the N specialized processors constitute the complete optimizer data. The value of N is an integer greater than or equal to 2.

[0138] In some embodiments, the optimizer data includes M optimizer parameters, M variances, and M momenta. The general processor may divide the M optimizer parameters, M variances, and M momenta into N parts. Alternatively, the general processor may divide the M optimizer parameters, M variances, and M momenta into N data sets. A dedicated processor in the group contains one data set. Each data set contains optimizer parameters, variances, and momentum. Different dedicated processors in the group contain different data sets. The N data sets contained by N dedicated processors constitute the complete optimizer data. The general processor configures the first data set to dedicated processor 1 of group 1, the second data set to dedicated processor 2 of group 1, and so on, and configures the Nth data set to dedicated processor N of group 1.

[0139] The dataset contains the same number of optimizer parameters, variances, and momenta. A dataset can contain one or more sets of optimizer parameters, variances, and momenta. For example, a dataset can contain one optimizer parameter, one variance, and one momenta. Another example can contain multiple optimizer parameters, multiple variances, and multiple momenta.

[0140] Different datasets can contain the same number of optimizer parameters. For example, each dataset contains M / N optimizer parameters, M / N variances, and M / N momentum.

[0141] Different datasets can also contain different numbers of optimizer parameters. For example, the first dataset contains 100 optimizer parameters, 100 variances, and 100 momentums. The second dataset contains 150 optimizer parameters, 150 variances, and 150 momentums.

[0142] If different datasets contain the same number of optimizer parameters, the different datasets also contain the same number of momentum and variance. If different datasets contain different numbers of optimizer parameters, the different datasets also contain different numbers of momentum and variance.

[0143] For example, a group includes two dedicated processors that divide the optimizer data into two data sets. The two dedicated processors in the group include the two data sets: the first dedicated processor in the group includes the first data set of the optimizer data, and the second dedicated processor in the group includes the second data set of the optimizer data. The two data sets included by the two dedicated processors constitute the complete initial optimizer data. The optimizer data includes 175 billion optimizer parameters, 175 billion variances, and 175 billion momenta. Each dedicated processor in the group includes 87.5 billion optimizer parameters, 87.5 billion variances, and 87.5 billion momenta. The first dedicated processor includes optimizer parameters, variances, and momenta labeled from 1 to 87.5 billion. The second dedicated processor includes optimizer parameters, variances, and momenta labeled from 87.6 billion to 175 billion.

[0144] Step 530: The dedicated processor initializes the optimizer data.

[0145] The dedicated processors in the group transmit the configured optimizer data to the dedicated processors of other groups, so that the dedicated processors of other groups contain the optimizer data, so that each group contains a complete copy of the optimizer data, and multiple groups contain multiple copies of the complete optimizer data, thus realizing global multiple copies of the optimizer data.

[0146] For example, dedicated processor 1 of group 1 transmits the first data set to dedicated processor 1 of group R, dedicated processor 2 of group 1 transmits the second data set to dedicated processor 2 of group R, and so on, dedicated processor N of group 1 transmits the Nth data set to dedicated processor N of group R.

[0147] For example, assuming R equals 2 and N equals 2, as shown in FIG6(a), the dedicated processors in the system are divided into two groups, each group containing two dedicated processors. The optimizer data includes a first data set and a second data set. Dedicated processor 1 in group 1 receives the first data set, and dedicated processor 2 in group 1 receives the second data set. Dedicated processor 1 in group 1 transmits the first data set to dedicated processor 1 in group 2. Dedicated processor 2 in group 1 transmits the second data set to dedicated processor 2 in group 2.

[0148] After the initialization process is completed, the dedicated processor trains the model. As shown in FIG5(b), the optimizer data update process. The method includes the following steps 540 to 560.

[0149] Step 540: The dedicated processor trains the model to obtain a gradient.

[0150] Dedicated processors within each group train the model based on the training data and generate gradients. For example, dedicated processors perform forward, activation, backward, gradient generation, and gradient accumulation operations.

[0151] In some embodiments, the optimizer data includes M optimizer parameters, M variances, and M momenta. A dedicated processor in each data parallel domain trains the model based on the training data to obtain M gradients.

[0152] Step 550: The dedicated processor transmits the gradient to a dedicated processor containing the same data set.

[0153] The dedicated processor transmits the gradient to the dedicated processor where the optimizer data having the same identifier as the gradient is located.

[0154] For example, dedicated processor 1 in group 1 and dedicated processor 1 in group 2 each include a first data set, which includes optimizer data from identifier 1 to identifier M / 2. Dedicated processor 2 in group 1 and dedicated processor 2 in group 2 each include a second data set, which includes optimizer data from identifier (M / 2)+1 to identifier M.

[0155] Specialized processor 1 in group 1 transmits the gradient from label 1 to label M / 2 to specialized processor 1 in group 2, and specialized processor 2 in group 1 transmits the gradient from label (M / 2)+1 to label M to specialized processor 2 in group 2.

[0156] For example, as shown in Figure 7, all cards in the computing system are divided into four data-parallel domains, each containing 250 cards. The optimizer data includes M optimizer parameters, M variances, and M momenta. All cards in the computing system are divided into two groups, and the optimizer data is divided into 250 data sets, with each card in each group containing one data set. Each data set includes M / 250 optimizer parameters, M / 250 variances, and M / 250 momenta. The gradients generated by the cards in the four data-parallel domains are transmitted to the cards in group 1 and group 2 containing the optimizer data with the same gradient identifier.

[0157] For example, as shown in FIG6(b), dedicated processor 1 in group 1 generates gradient 1 and gradient 2, and dedicated processor 2 in group 1 generates gradient 3 and gradient 4. Dedicated processor 1 in group 1 transmits gradient 1 and gradient 2 to dedicated processor 1 in group 2. Dedicated processor 2 in group 1 transmits gradient 3 and gradient 4 to dedicated processor 1 in group 1 and dedicated processor 1 in group 2.

[0158] Dedicated processor 1 in group 2 generates gradient 5 and gradient 6, and dedicated processor 2 in group 2 generates gradient 7 and gradient 8. Dedicated processor 1 in group 2 transmits gradient 5 and gradient 6 to dedicated processor 2 in group 1 and dedicated processor 2 in group 2. Dedicated processor 2 in group 2 transmits gradient 7 and gradient 8 to dedicated processor 2 in group 1.

[0159] In some embodiments, if the first dataset contains optimizer parameters, variances, and momentum corresponding to multiple identifiers, the gradients corresponding to the multiple identifiers are transmitted to the dedicated processors corresponding to the datasets with the same identifiers. For example, if the first dataset contains optimizer parameters, variances, and momentum corresponding to identifiers 1 through 3, and dedicated processor 1 in group 1 and dedicated processor 1 in group 2 both contain the first dataset, dedicated processor 1 in group 1 will transmit the gradients corresponding to identifiers 1 through 3 to dedicated processor 1 in group 2.

[0160] Step 560: The dedicated processor updates the optimizer data according to the gradient.

[0161] Each dedicated processor updates the data set in the optimizer data contained in the dedicated processor according to the gradient to obtain an updated data set.

[0162] The dedicated processor 1 in group 1 and the dedicated processor 1 in group R update the first data set according to the gradient. Similarly, the dedicated processor N in group 1 and the dedicated processor N in group R update the Nth data set according to the gradient.

[0163] For example, dedicated processor 1 in group 1 and dedicated processor 1 in group 2 both update the optimizer parameters, variance, and momentum (parameter momentum variance, PMV) of identifier 1 according to the gradient of identifier 1. Similarly, dedicated processor 1 in group 1 and dedicated processor 1 in group 2 both update the optimizer parameters, variance, and momentum of identifier M / 2 according to the gradient of identifier M / 2 to obtain updated optimizer parameters, updated variance, and updated momentum.

[0164] The dedicated processor 2 in group 1 and the dedicated processor 2 in group 2 both update the optimizer parameters, variance and momentum of the identifier (M / 2)+1 according to the gradient of the identifier (M / 2)+1. Similarly, the dedicated processor 2 in group 1 and the dedicated processor 2 in group 2 both update the optimizer parameters, variance and momentum of the identifier M according to the gradient of the identifier M to obtain updated optimizer parameters, updated variance and updated momentum.

[0165] Regarding the methods for updating optimizer parameters, variance, and momentum, reference can be made to traditional methods without limitation.

[0166] The optimizer data update scheme provided by this application includes ensuring that the optimizer data between multiple groups is the same during initialization. During the model training process, gradients are transmitted between multiple groups. Then, based on the same optimizer data (previous PMV), the same input (gradient), and the same optimizer algorithm, the dedicated processor makes the updated optimizer data between multiple groups the same, that is, the update results (this PMV) are the same. Therefore, because the amount of gradient data is much smaller than the amount of optimizer data, for example, the amount of gradient data is 1 / 6 of the amount of optimizer data, only gradients are transmitted between dedicated processors, reducing the amount of data transmitted and bandwidth. The dedicated processor updates the optimizer data based on the gradient, ensuring that the updated optimizer data between multiple groups is the same, realizing "calculation instead of transmission", and ensuring the integrity of the optimizer data at all times.

[0167] After the optimizer data is updated, the dedicated processor updates the model parameters. As shown in FIG5 (c), the process of updating the model parameters. The method includes the following steps 570 to 580.

[0168] Step 570: The dedicated processor converts the optimizer parameters into model parameters.

[0169] A dedicated processor converts single-precision optimizer parameters to half-precision model parameters, for example, converting optimizer parameters fp32 to model parameters fp16.

[0170] Step 580: The dedicated processor transmits the model parameters to a dedicated processor for training the model parameters.

[0171] After the dedicated processors convert the single-precision optimizer parameters to half-precision model parameters, they perform an AllGather operation. For example, each dedicated processor in R groups contains model parameter P1. Dedicated processor 1 in group 1 transmits model parameter P1 to each dedicated processor in the R groups. Each dedicated processor in R groups contains model parameter P2. Dedicated processor 2 in group 1 transmits model parameter P2 to each dedicated processor in the R groups.

[0172] For example, as shown in (c) of Figure 6 , dedicated processors 1 and 2 in group 1, as well as dedicated processors 1 and 2 in group 2, each contain model parameters P1 and P2. Dedicated processor 1 in group 1 converts the single-precision optimizer parameters into half-precision model parameters P1, and transmits the model parameters P1 to dedicated processors 1 and 2 in group 1, as well as dedicated processors 1 and 2 in group 2. Dedicated processor 2 in group 1 converts the single-precision optimizer parameters into half-precision model parameters P2, and transmits the model parameters P2 to dedicated processors 1 and 2 in group 1, as well as dedicated processors 1 and 2 in group 2.

[0173] This ensures that the optimizer data and model parameters are consistent, ensuring that the model is trained correctly.

[0174] In other embodiments, after a dedicated processor in the system fails, the non-faulty dedicated processor can persist the dying checkpoint without saving the model parameters. As shown in FIG8 , the present application further includes the following steps 590 to 5130 .

[0175] Assume that dedicated processor 1 in group 1 fails, and the non-faulty dedicated processors detect the presence of an abnormal dedicated processor in the computing system. The general-purpose processor instructs the persistence of the datasets contained by multiple non-faulty dedicated processors. The datasets contained by the multiple non-faulty dedicated processors constitute updated optimizer data. The multiple non-faulty dedicated processors include a first dedicated processor, and the updated optimizer data includes the updated first dataset. The datasets contained by the multiple non-faulty dedicated processors also include the updated first dataset.

[0176] Step 590: The non-faulty dedicated processor sends an exception message to the general processor.

[0177] If a dedicated processor in a computing system fails, other dedicated processors that are not faulty may not be able to receive data sent by the faulty dedicated processor. The remaining dedicated processors then sense the presence of the faulty dedicated processor in the computing system and send an exception message to the general processor. The exception message indicates the presence of the faulty dedicated processor in the system.

[0178] Step 5100: The general processor instructs the non-faulty dedicated processor to report integrity information.

[0179] The general processor sends a reporting command to the non-faulty dedicated processor, and the non-faulty dedicated processor reports the integrity information to the general processor. The reporting command is used to instruct the non-faulty dedicated processor to report the integrity information.

[0180] The integrity information is used to indicate the integrity of the data set contained in the non-faulty dedicated processor. For example, the integrity information includes the identification and correctness information of the data in the data set. The integrity information also includes the number of iterations of the data in the data set contained in the non-faulty dedicated processor.

[0181] Step 5110: Determine, based on the integrity information sent by the non-faulty dedicated processors in the computing system, that the data sets contained in the plurality of non-faulty dedicated processors constitute updated optimizer data.

[0182] Because the optimizer data contains at least two copies of the dataset stored in the computing system, the integrity information received by the general-purpose processor may contain the same integrity information, that is, at least two dedicated processors sent the same identifier to the general-purpose processor. It should be noted that if the optimizer data contains two copies of the dataset stored in the system and one dedicated processor fails, the general-purpose processor will receive the integrity information for the dataset contained in the failed dedicated processor from the non-faulty dedicated processor. The general-purpose processor may also receive the integrity information for the dataset contained in two copies of the non-faulty dedicated processor from other non-faulty dedicated processors.

[0183] The general processor filters multiple continuous identifiers from the received identifiers, and the data indicated by the multiple continuous identifiers constitute the updated optimizer data.

[0184] The general purpose processor determines the non-faulty dedicated processors where the data indicated by the multiple consecutive identifiers are located as the multiple non-faulty dedicated processors constituting updated optimizer data. The updated optimizer data includes updated optimizer parameters, updated variance, and updated momentum.

[0185] Understandably, the data contained in multiple non-faulty dedicated processors constitute the complete optimizer data. That is, the general processor determines the dedicated processors that can constitute the complete optimizer data based on the multiple consecutive identifiers contained in the integrity information.

[0186] Optionally, the general processor can also determine that the data indicated by multiple consecutive identifiers are correct based on the correct information contained in the integrity information, and then determine the non-faulty dedicated processors where the data indicated by multiple consecutive identifiers are located as the multiple non-faulty dedicated processors that constitute the updated optimizer data.

[0187] Optionally, the integrity information further includes an iteration count. The general-purpose processor determines, based on the iteration count included in the integrity information, that the data indicated by the multiple consecutive identifiers have the same iteration count, and then determines the non-faulty dedicated processors containing the data indicated by the multiple consecutive identifiers as the non-faulty dedicated processors constituting the updated optimizer data.

[0188] It can be understood that the multiple identifiers constituting the complete optimizer data determined by the general processor are continuous, the number of iterations of the data indicated by the multiple identifiers is the same, and the data are all correct.

[0189] For example, the optimizer data includes a first data set and a second data set. The dedicated processor 1 in group 1 and the dedicated processor 1 in group 2 contain the first data set. The dedicated processor 2 in group 1 and the dedicated processor 2 in group 2 contain the second data set. If the dedicated processor 1 in group 1 fails, the dedicated processor 1 in group 2 reports the integrity information of the first data set to the general processor, and the dedicated processor 2 in group 1 and the dedicated processor 2 in group 2 report the integrity information of the second data set to the general processor. The general processor determines that the first data set contained in the dedicated processor 1 in group 2 and the second data set contained in the dedicated processor 2 in group 1 constitute complete optimizer data. The general processor then instructs the dedicated processor 1 in group 2 to persist the first data set, and instructs the dedicated processor 2 in group 1 to persist the second data set.

[0190] Step 5120: The general processor sends a persistence command to multiple non-faulty dedicated processors.

[0191] The persist command is used to instruct the persistence of the data sets contained in multiple non-faulty dedicated processors. The general processor instructs multiple non-faulty dedicated processors to persist the data sets contained in the optimizer data.

[0192] Step 5130: The dedicated processor persists the data.

[0193] The dedicated processor receives the persist command sent by the general-purpose processor and persists the optimizer data set contained in the dedicated processor to the storage system. For example, if the optimizer data includes a first data set and a second data set, the first data set contained in the optimizer data contained in dedicated processor 1 is stored in the storage system, and the second data set contained in the optimizer data contained in dedicated processor 2 is also stored in the storage system. The storage system can store the optimizer data in the form of files.

[0194] It can be understood that the storage system stores the final checkpoint, that is, stores the complete optimizer data in the system when the dedicated processor fails, and the optimizer data includes the updated optimizer parameters, the updated variance and the updated momentum.

[0195] Therefore, due to the multiple copies of optimizer data in the global system, when a dedicated processor fails, since the system contains multiple copies of optimizer data, the complete optimizer data can be composed through the data sets contained in the non-faulty dedicated processors, ensuring the integrity of the optimizer data when the dedicated processor fails, improving the reliability of the optimizer data, and ensuring the accuracy of the dying checkpoint, so as to facilitate the resumption of model training based on the dying checkpoint and reduce training losses.

[0196] Optionally, the above embodiment is described by taking a general-purpose processor executing steps 5100 to 5130 as an example. In some embodiments, the general-purpose processor may also designate a dedicated processor to execute steps 5100 to 5130.

[0197] In other embodiments, after a dedicated processor in the system fails, the dedicated processor can resume model training based on a dying checkpoint. Figure 9 is a flow chart of a data processing method provided by the present application. Here, the recovery model training is mainly described. As shown in (a) of Figure 9, the initialization process before the dedicated processor trains the model, the method includes the following steps 910 to 930. After the optimizer data is updated, the dedicated processor updates the model parameters. As shown in (b) of Figure 9, the process of updating the model parameters. The method includes the following steps 940 to 960.

[0198] Step 910: The general processor obtains a request.

[0199] The general processor receives a request from the client, which is used to instruct the resumption of model training. The request may include a model identifier, so that the general processor can identify the model based on the model identifier and instruct the dedicated processor to load the model indicated by the model identifier from the storage system.

[0200] The general-purpose processor can obtain optimizer data from the storage system. When the general-purpose processor instructs the dedicated processor to resume training, that is, when the dedicated processor performs breakpoint-resumed training, the general-purpose processor obtains updated optimizer data from the storage system. The updated optimizer data includes updated values ​​for optimizer parameters, momentum, and variance. The updated optimizer data can be the optimizer data included in the final checkpoint.

[0201] Step 920: The general processor initializes the updated optimizer data.

[0202] The general-purpose processor can group the dedicated processors in the system to obtain multiple groups. The general-purpose processor selects one of the multiple groups and divides the updated optimizer data according to the number of dedicated processors in the group. Each dedicated processor in the group contains a portion of the updated optimizer data. The updated optimizer data contained in at least two dedicated processors belonging to different groups is the same. Each group contains a complete copy of the updated optimizer data, and multiple groups contain multiple copies of the complete updated optimizer data, thereby realizing global multiple copies of the updated optimizer data. For an explanation of the general-purpose processor initializing the dedicated processor, please refer to the description of step 520 above.

[0203] It should be noted that the dedicated processors included in each group after this grouping can be the same as or different from the dedicated processors included in each group after the previous grouping, and there is no limitation. After the faulty dedicated processor is restored, the current grouping can also include the restored dedicated processor.

[0204] Step 930: The dedicated processor initializes the updated optimizer data.

[0205] The dedicated processors in the group transmit the configured optimizer data to the dedicated processors in other groups, so that the dedicated processors in other groups contain the updated optimizer data. Each group contains a complete copy of the updated optimizer data. Multiple groups contain multiple copies of the complete updated optimizer data, thus realizing global multiple copies of the optimizer data.

[0206] It should be noted that step 930 is an optional step. In some embodiments, the general processor can also obtain the updated optimizer data from the storage system again, divide the updated optimizer data into N parts according to the number N of dedicated processors in another group, and configure 1 / N data in the updated optimizer data for a dedicated processor in another group.

[0207] Step 940: The dedicated processor converts the optimizer parameters into model parameters.

[0208] A dedicated processor converts single-precision optimizer parameters to half-precision model parameters, for example, converting optimizer parameters fp32 to model parameters fp16.

[0209] Step 950: The dedicated processor transmits the model parameters to a dedicated processor for training the model parameters.

[0210] The dedicated processor transmits the gradient to the dedicated processor where the optimizer data with the same identifier as the gradient is located. For explanation of step 950 , reference may be made to the explanation of step 580 above.

[0211] Step 960: The dedicated processor trains the model.

[0212] Since the dying checkpoint only saves the optimizer data and not the model parameters, when resuming model training, the model parameters are calculated based on the optimizer data in the storage system and then transmitted to the dedicated processor used to train the model parameters, thereby ensuring that the optimizer data and model parameters at the dying checkpoint are consistent. This overcomes the problem of inconsistency between the optimizer data and model parameters due to the failure of the AllGather operation when the dedicated processor fails. In addition, the method provided in this application can save the checkpoint at the time of the failure and resume model training based on the checkpoint at the time of the failure, thereby effectively reducing training losses and improving the efficiency of model training.

[0213] Figure 10 is a schematic diagram comparing optimizer data storage provided by this application. As shown in Figure 10 (a), all cards in the computing system are divided into four data parallel domains, each containing four cards. The optimizer data includes M optimizer parameters, M variances, and M momenta. The optimizer data is divided into 16 parts, with each card containing 1 / 16 of the data.

[0214] As shown in Figure 10(b), all cards in the computing system are divided into two groups, and the optimizer data is divided into eight data sets, with each card in each group containing one data set. Each data set includes M / 8 optimizer parameters, M / 8 variances, and M / 8 momentum. Gradients generated by the cards in the four data-parallel domains are transmitted to the cards in group 1 and group 2, which contain the optimizer data with the same gradient identifier.

[0215] Therefore, if the optimizer data is replicated brute force, the data transfer volume is 12MB, which is the total amount of data for the entire optimizer. The method provided in this application reduces the data transfer volume to 2MB, which is the amount of data for the gradient. Deploying at least two copies of the optimizer data globally allows for multiple copies of the optimizer data globally.

[0216] Figure 11 is a schematic diagram of a data processing device 1100 provided by the present application. During the model training process, the optimizer multi-copy device 1101 is used to perform calculation-by-transfer to update the optimizer data of multiple copies and ensure the integrity of the optimizer data at all times.

[0217] When a system anomaly occurs, the process of saving the final checkpoint includes a training anomaly acquisition device 1102, an optimizer data global storage device 1103, and a distributed job cluster 1104. The training anomaly acquisition device 1102 is used to detect system anomalies. The optimizer data global storage device 1103 is used to store the optimizer data set.

[0218] The work cluster 1104 is used to perform health status detection of workers, perform integrity checks, and issue persistence commands for end-of-life checkpoints.

[0219] 12 is a schematic diagram of another data processing device provided by the present application. The data processing device 1200 includes a dying checkpoint recovery device, which includes an optimizer state loading module 1201, a parameter generator 1202, and a parameter propagator 1203.

[0220] The optimizer state loading module 1201 is used to read the optimizer data from the dying checkpoint file and synchronize the copy of the optimizer data.

[0221] Parameter generator 1202 is used to generate model parameters based on optimizer parameters. Since optimizer data contains optimizer parameters, but the optimizer parameters and model parameters are of different types, for example, if the optimizer parameters are fp32 and the model parameters are fp16 during the forward and reverse operations, if the optimizer parameters are inconsistent with the model parameters, the fp32 will be converted to fp16.

[0222] The parameter propagator 1203 is used to synchronize the model parameters to all corresponding training workers through communication operations. Once the model parameters are propagated, the training iteration can begin.

[0223] The functions of the various devices in the data processing devices shown in FIG. 11 and FIG. 12 may be explained with reference to the description of the above embodiments.

[0224] It is understood that in order to implement the functions in the above embodiments, the computer device includes hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a manner driven by computer software depends on the specific application scenario and design constraints of the technical solution.

[0225] The data processing method provided by the present application is described in detail above in conjunction with Figures 1 to 13. The device provided by the present application will be described below in conjunction with Figure 13. These devices can be used to implement the functions of the dedicated processor or general-purpose processor in the above method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments. In this embodiment, the device can be a dedicated processor or general-purpose processor as shown in Figures 5, 8, and 9, or a module (such as a chip) applied to a computer device.

[0226] As shown in (a) of FIG13 , the data processing device 1300 includes a communication module 1301 , a training module 1302 , an optimizer data updating module 1303 and a storage module 1304 .

[0227] The data processing device 1300 is used to implement the functions of the dedicated processor in the method embodiments shown in FIG. 5 , FIG. 8 , and FIG. 9 .

[0228] The training module 1302 is used to train the artificial intelligence model to obtain the first gradient. For example, the training module 1302 is used to execute step 960 in FIG. 9 and step 540 in FIG. 5 .

[0229] The optimizer data updating module 1303 is configured to update the first data set according to the first gradient to obtain an updated first data set. For example, the optimizer data updating module 1303 is configured to execute step 560 in FIG5 .

[0230] Optionally, the communication module 1301 is configured to transmit the first gradient. For example, the communication module 1301 is configured to execute step 550 in FIG5 .

[0231] Optionally, the communication module 1301 is further configured to obtain an initial value of a first data set configured for the general processor. For example, the communication module 1301 is configured to execute step 510 in FIG5 .

[0232] Optionally, the optimizer data updating module 1303 is further configured to convert the updated optimizer parameters in the first data set into model parameters. For example, the optimizer data updating module 1303 is configured to execute step 570 in FIG5 and step 940 in FIG9.

[0233] Optionally, the communication module 1301 is further configured to transmit the model parameters to a dedicated processor in the computing system for training the model parameters. For example, the communication module 1301 is configured to execute step 580 in FIG5 and step 950 in FIG9.

[0234] Optionally, the optimizer data update module 1303 is further configured to persist the included data set. For example, the optimizer data update module 1303 is configured to execute step 5130 in FIG8 .

[0235] Optionally, the communication module 1301 is further configured to obtain an initial value of the first data set. For example, the communication module 1301 is configured to execute step 530 in FIG5 .

[0236] The storage module 1304 is used to store the data set, gradient, optimization algorithm, etc. in the optimizer data to facilitate model training.

[0237] As shown in (b) of FIG. 13 , the data processing device 1300 is used to implement the functions of the general-purpose processor in the method embodiment shown in FIG. 8 .

[0238] The data processing device 1300 includes a communication module 1301 , a fault handling module 1305 , an optimizer data updating module 1303 and a storage module 1304 .

[0239] Fault handling module 1305 is configured to instruct, when a dedicated processor in the system fails, to persist data sets included in multiple non-faulty dedicated processors, wherein the data sets included in the multiple non-faulty dedicated processors constitute updated optimizer data, the multiple non-faulty dedicated processors include the first dedicated processor, and the updated optimizer data includes the updated first data set. For example, fault handling module 1305 is configured to execute step 5110 in FIG. 8 .

[0240] Optionally, the communication module 1301 is configured to obtain integrity information sent by a non-faulty dedicated processor in the computing system, where the integrity information indicates the integrity of the data set contained in the non-faulty dedicated processor. For example, the communication module 1301 is configured to execute step 5100 in FIG8 .

[0241] Optionally, when the fault handling module 1305 instructs to persist the data sets contained in multiple non-faulty dedicated processors, it is specifically used to: determine that the data sets contained in multiple non-faulty dedicated processors constitute the updated optimizer data based on the integrity information sent by the non-faulty dedicated processors in the computing system.

[0242] Optionally, the communication module 1301 is further configured to send a persistence command to the plurality of non-faulty dedicated processors, the persistence command being used to instruct persistence of the data sets contained in the plurality of non-faulty dedicated processors. For example, the communication module 1301 is configured to execute step 5120 in FIG8 .

[0243] Optionally, the integrity information includes the identification and correctness information of the data in the data set; when the fault processing module 1305 determines that the data sets contained in multiple non-faulty dedicated processors constitute the updated optimizer data based on the integrity information sent by the non-faulty dedicated processors in the computing system, it is specifically used to: determine multiple consecutive identifiers based on the identifiers contained in the integrity information sent by the non-faulty dedicated processors in the computing system, and the data indicated by the multiple consecutive identifiers constitute the updated optimizer data; determine that the data indicated by the multiple consecutive identifiers are correct based on the correctness information contained in the integrity information; and determine the non-faulty dedicated processors where the data indicated by the multiple consecutive identifiers are located as the multiple non-faulty dedicated processors constituting the updated optimizer data.

[0244] Optionally, the integrity information further includes the number of iterations of data in the data set included in the non-faulty dedicated processor; the fault processing module 1305 is further used to determine, based on the number of iterations included in the integrity information, whether the number of iterations of data indicated by multiple consecutive identifiers is the same.

[0245] Optionally, the communication module 1301 is further used to obtain updated optimizer data when restarting model training; the optimizer data update module 1303 is used to configure the updated first data set to the first dedicated processor.

[0246] Optionally, the communication module 1301 is further configured to configure the updated first data set to the third dedicated processor.

[0247] The storage module 1304 is used to store integrity information, data sets in the optimizer data, etc.

[0248] It should be understood that the data processing device 1300 of the embodiment of the present application can be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), wherein the PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. Alternatively, when the methods shown in FIG. 5 , FIG. 8 , and FIG. 9 are implemented by software, their respective modules can also be software modules, and the computation graph optimization device 900 and its respective modules can also be software modules.

[0249] According to the data processing device 1300 of the embodiment of the present application, it can correspond to executing the method described in the embodiment of the present application, and the above-mentioned and other operations and / or functions of each unit in the data processing device 1300 are respectively for realizing the corresponding processes of each method in Figures 5, 8, and 9. For the sake of brevity, they will not be repeated here.

[0250] FIG14 is a schematic diagram of the structure of a computer device 1400 provided in this application. As shown in FIG14 , computer device 1400 includes a processor 1410, a bus 1420, a memory 1430, a communication interface 1440, a memory 1450 (also referred to as a main memory unit), and a processor 1460. Processor 1410, processor 1460, memory 1430, memory 1450, and communication interface 1440 are connected via bus 1420.

[0251] It should be understood that in this embodiment, the processor 1410 may be a CPU, but may also be other general-purpose processors, digital signal processors (DSP), ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0252] Computer device 1400 may also include a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application. For example, processor 1460 may be a GPU or an NPU.

[0253] The communication interface 1440 is used to implement communication between the computer device 1400 and external devices or components.

[0254] In this application, when the computer device 1400 is used to implement the functions of the dedicated processor shown in Figures 5, 8, and 9, the communication interface 1440 is used to transmit gradients, initial values, etc., so that the processor 1460 can be used to train the model, update the optimizer data, and persist the final checkpoint.

[0255] When the computer device 1400 is used to implement the functions of the general processor shown in Figures 5, 8, and 9, the communication interface 1440 is used to transmit integrity information, etc., so that the processor 1410 is used to instruct the persistence of data sets included in multiple non-faulty dedicated processors.

[0256] The bus 1420 may include a path for transmitting information between the above-mentioned components (such as the processor 1410, the memory 1450, and the storage 1430). In addition to the data bus, the bus 1420 may also include a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus 1420 in the figure. The bus 1420 may be a Peripheral Component Interconnect Express (PCIe) bus, an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus 1420 can be divided into an address bus, a data bus, a control bus, etc.

[0257] As an example, computer device 1400 may include multiple processors. The processor may be a multi-core (multi-CPU) processor. A processor herein may refer to one or more devices, circuits, and / or computing units for processing data (e.g., computer program instructions).

[0258] It is worth noting that FIG14 only uses a computer device 1400 including one processor 1410 and one memory 1430 as an example. Here, the processor 1410 and the memory 1430 are respectively used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined based on business requirements. For example, the computer device 1400 may include multiple GPUs or NPUs.

[0259] Memory 1450 may be a volatile memory pool or a nonvolatile memory pool, or may include both volatile and nonvolatile memory. Nonvolatile memory may be read-only memory (ROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). Memory 1450 is used to store optimizer data, gradients, and the like.

[0260] The memory 1430 may correspond to a storage medium used to store optimizer data and other information in the above method embodiment, for example, a disk such as a mechanical hard disk or a solid-state drive.

[0261] The computer device 1400 may be a general-purpose device or a dedicated device. For example, the computer device 1400 may be a server or other device with computing capabilities.

[0262] It should be understood that the computer device 1400 according to this embodiment may correspond to the data processing device 1300 in this embodiment, and may correspond to executing the corresponding subject in any method in Figures 5, 8, and 9, and the above-mentioned and other operations and / or functions of each module in the data processing device 1300 are respectively for realizing the corresponding processes of each method in Figures 5, 8, and 9. For the sake of brevity, they will not be repeated here.

[0263] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a computing device. Of course, the processor and storage medium can also exist as discrete components in a computing device.

[0264] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD). The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A data processing method, characterized in that: Applied to a computing system, the computing system includes a general-purpose processor and multiple special-purpose processors, the multiple special-purpose processors are used to train an artificial intelligence model, the multiple special-purpose processors include at least two optimizer data, the optimizer data includes a first data set, the multiple special-purpose processors include a first special-purpose processor and a second special-purpose processor, the first special-purpose processor and the second special-purpose processor both include a first data set, the method includes: The first dedicated processor trains the artificial intelligence model to obtain a first gradient; The first dedicated processor updates the first data set according to the first gradient to obtain an updated first data set; The second dedicated processor updates the first data set according to the first gradient to obtain the updated first data set; When the second dedicated processor fails, the general processor instructs to persist the data sets contained in multiple non-faulty dedicated processors, the data sets contained in the multiple non-faulty dedicated processors constitute the updated optimizer data, the multiple non-faulty dedicated processors include the first dedicated processor, and the updated optimizer data includes the updated first data set.

2. The method according to claim 1, characterized in that The general processor instructs the persistence of a data set included in a plurality of non-faulty dedicated processors, including: Acquire integrity information sent by a non-faulty dedicated processor in the computing system, the integrity information being used to indicate integrity of a data set contained in the non-faulty dedicated processor; determining, based on integrity information sent by non-faulty dedicated processors in the computing system, that data sets included in the plurality of non-faulty dedicated processors constitute the updated optimizer data; A persistence command is sent to the plurality of non-faulty dedicated processors, where the persistence command is used to instruct persistence of the data sets included in the plurality of non-faulty dedicated processors.

3. The method according to claim 2, characterized in that The integrity information includes the identification and correct information of the data in the data set; Determining, based on the integrity information sent by the non-faulty dedicated processors in the computing system, that the data sets included in the plurality of non-faulty dedicated processors constitute the updated optimizer data, comprises: Determine a plurality of continuous identifiers according to identifiers included in the integrity information sent by the non-faulty dedicated processor in the computing system, wherein the data indicated by the plurality of continuous identifiers constitute the updated optimizer data; Determining, based on correct information included in the integrity information, that the data indicated by the multiple consecutive identifiers are correct; The non-faulty dedicated processors where the data indicated by the plurality of consecutive identifiers are located are determined as the plurality of non-faulty dedicated processors constituting the updated optimizer data.

4. The method according to claim 3, characterized in that The integrity information also includes the number of iterations of data in the data set included in the non-faulty dedicated processor; The method further comprises: It is determined, according to the number of iterations included in the integrity information, that the number of data iterations indicated by the multiple consecutive identifiers is the same.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The first dedicated processor obtains an initial value of the first data set configured by the general processor; The second dedicated processor obtains an initial value of the first data set sent by the first dedicated processor.

6. The method according to any one of claims 1 to 5, characterized in that The optimizer data further includes a second data set, and the plurality of special-purpose processors include a third special-purpose processor and a fourth special-purpose processor, and the third special-purpose processor and the fourth special-purpose processor both contain the second data set.

7. The method according to any one of claims 1 to 6, characterized in that After the general processor instructs to persist the data sets included in the plurality of non-faulty dedicated processors, the method further comprises: When restarting model training, the general processor obtains the updated optimizer data; The general processor configures the updated first data set to the first dedicated processor; The general processor configures the updated first data set to the third special processor.

8. The method according to claim 7, characterized in that The method further comprises: The first dedicated processor converts the updated first data center optimizer parameters into model parameters; The first dedicated processor transmits the model parameters to a dedicated processor in the computing system for training the model parameters.

9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: The first special purpose processor transmits the first gradient to the second special purpose processor.

10. The method according to any one of claims 1 to 9, characterized in that The optimizer data includes optimizer parameters and optimizer state, and the optimizer state includes variance and momentum.

11. A computer device, characterized in that: The computer device includes a memory and multiple processors, the memory is used to store a set of computer instructions; when the processor executes the set of computer instructions, the multiple processors jointly execute the operating steps of any one of the methods described in claims 1-10.

Citation Information

Patent Citations

  • Data processing method and computer equipment

    CN120123142A

  • Training device and method of neural network model and related equipment

    CN113705801A

  • Multi-dimensional parallel processing method, system and device based on artificial intelligence, and readable storage medium

    CN114035936A

  • Distributed weight update for back propagation of neural networks

    CN114631102A

  • Server and data center

    CN115794381A

Cited By

  • Text classification method and device based on large model and task execution method and device based on large model

    CN121233773A