Deep learning model training method and deep learning model training system
During the distributed training process of deep learning models, the target storage parameters are stored in real time based on the model specification information and the training strategy, and the model parameters are stored in real time, which solves the progress loss and performance overhead problems caused by training exceptions, and achieves high stability and high efficiency training.
Patent Information
- Application Number
- CN202311636365.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-11-30
AI Technical Summary
During distributed training, training exceptions often lead to training interruptions, resulting in large training progress losses. The storage of updated model parameters after each iteration increases performance overhead and reduces training efficiency.
During the calculation process of distributed training, the model parameters of the deep learning model are stored according to the target storage parameters. The target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, real-time storage of model parameters is realized, and the stored procedures are overlapped with the training process, saving performance overhead.
It realizes high stability and efficiency of deep learning model training, avoids progress losses caused by training exceptions, reduces performance overhead, and improves the fault tolerance and efficiency of training.
Smart Images

Figure CN117669700B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of deep learning technology, and in particular, to a deep learning model training method and a deep learning model training system. Background Art
[0002] With the development of deep learning technology, large-scale deep learning models represented by large language models have been widely used in tasks in different scenarios.
[0003] At present, efficient training of deep learning models through distributed training has become the mainstream method of model training. However, in the distributed training process, training anomalies are inevitable, which often lead to training interruptions, and the updated model parameters are not stored, resulting in a large loss of training progress and insufficient stability of deep learning model training. In the distributed training process, the updated model parameters are stored after each iteration, which increases performance overhead and reduces training efficiency. Therefore, a highly stable and efficient deep learning model training method is urgently needed. Summary of the invention
[0004] In view of this, the embodiments of this specification provide a deep learning model training method. One or more embodiments of this specification also involve another deep learning model training method, a deep learning model training system, a deep learning model training device, another deep learning model training device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a deep learning model training method is provided, including:
[0006] Obtain an initial deep learning model and sample dataset;
[0007] According to the preset distributed training strategy, the deep learning model is distributedly trained based on the sample data set, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.
[0008] According to a second aspect of an embodiment of this specification, another deep learning model training method is provided, which is applied to a cloud-side device, where the cloud-side device includes a plurality of distributed nodes and a storage medium; the method includes:
[0009] Obtain an initial deep learning model and sample dataset;
[0010] According to a preset distributed training strategy, multiple distributed nodes are called to perform distributed training on the deep learning model based on a sample data set, and in a process of calculating adjustment parameters of the distributed training, model parameters of the deep learning model are stored in a storage medium according to target storage parameters, wherein the target storage parameters are determined based on model specification information of the deep learning model and the preset distributed training strategy;
[0011] When an abnormality in deep learning model training is identified, multiple distributed nodes are triggered to stop distributed training of the deep learning model, and target model parameters currently stored in the storage medium are determined;
[0012] Upon receiving a request to resume training, obtaining target model parameters from a storage medium;
[0013] Call multiple distributed nodes to perform distributed training of deep learning models based on target model parameter recovery.
[0014] According to a third aspect of an embodiment of this specification, a deep learning model training system is provided, the system comprising a management and control unit and a plurality of distributed nodes, the plurality of distributed nodes comprising a first distributed node, the first distributed node being any one of the plurality of distributed nodes;
[0015] The control unit is used to obtain the initial deep learning model and sample data set, build multiple distributed data based on the deep learning model and sample data set according to the preset distributed training strategy, and distribute the multiple distributed data to each distributed node;
[0016] The first distributed node is used to perform distributed training on the deep learning model based on the sample data set; and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored therein according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.
[0017] According to a fourth aspect of the embodiments of this specification, a deep learning model training device is provided, including:
[0018] A first acquisition module is configured to acquire an initial deep learning model and a sample data set;
[0019] The first training module is configured to perform distributed training on the deep learning model based on the sample data set according to a preset distributed training strategy, and store the model parameters of the deep learning model according to the target storage parameters during the adjustment parameter calculation process of the distributed training, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.
[0020] According to a fifth aspect of an embodiment of this specification, another deep learning model training device is provided, which is applied to a cloud-side device, wherein the cloud-side device includes a plurality of distributed nodes and a storage medium; the device includes:
[0021] A second acquisition module is configured to acquire an initial deep learning model and a sample data set;
[0022] A second training module is configured to call multiple distributed nodes according to a preset distributed training strategy, perform distributed training on the deep learning model based on the sample data set, and store the model parameters of the deep learning model to a storage medium according to a target storage parameter during the calculation of adjustment parameters of the distributed training, wherein the target storage parameter is determined based on model specification information of the deep learning model and the preset distributed training strategy;
[0023] A stop module is configured to trigger multiple distributed nodes to stop the distributed training of the deep learning model when an abnormality in the deep learning model training is identified, and determine the target model parameters currently stored in the storage medium;
[0024] A parameter acquisition module is configured to acquire target model parameters from a storage medium when receiving a request to resume training;
[0025] The recovery module is configured to call multiple distributed nodes to perform distributed training on the deep learning model based on target model parameter recovery.
[0026] According to a sixth aspect of an embodiment of this specification, a computing device is provided, including:
[0027] Memory and processor;
[0028] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above method are implemented.
[0029] According to a seventh aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and the instructions implement the steps of the above method when executed by a processor.
[0030] According to an eighth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.
[0031] In one embodiment of the present specification, an initial deep learning model and a sample data set are obtained; distributed training is performed on the deep learning model based on the sample data set according to a preset distributed training strategy, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy. The target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, and the iterative law of distributed training is fully considered. In the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, and the storage process of the model parameters and the adjustment parameter calculation process in the distributed training are overlapped, which fully saves performance overhead, and completes the real-time storage of the deep learning model parameters in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of a deep learning model training method provided by an embodiment of this specification;
[0033] Figure 2 It is a system architecture diagram of a deep learning model training method provided by an embodiment of this specification;
[0034] Figure 3 It is a flowchart of a deep learning model training method provided by an embodiment of this specification;
[0035] Figure 4 It is a schematic diagram of a deep learning model training method provided by an embodiment of this specification;
[0036] Figure 5 is a flowchart of another deep learning model training method provided by an embodiment of this specification;
[0037] Figure 6 It is a structural diagram of a deep learning model training system provided by an embodiment of this specification;
[0038] Figure 7 It is a structural diagram of a deep learning model training device provided by an embodiment of this specification;
[0039] Figure 8 It is a structural diagram of another deep learning model training device provided by an embodiment of this specification;
[0040] Fig. 9 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0041] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0042] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0043] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0044] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0045] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, which usually contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than 10 trillion model parameters. A large model can also be called a foundation model / foundation model. A large model is pre-trained with large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (Large Language Model, LLM for short), a multi-modal pre-training model, etc.
[0046] When the big model is used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. The big model can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the big model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0047] First, the terms involved in one or more embodiments of this specification are explained.
[0048] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. Its purpose is to enable computers to understand and use human language to perform useful tasks. Natural language processing is mainly used in machine translation, speech recognition, text analysis, text question answering and other fields.
[0049] Hyperparameters: are fixed parameters calculated during the model training process. They can be understood as a strategy for updating model parameters and are used to control the update process of model parameters.
[0050] Gradient weight: refers to the relationship between the gradient in the model and the model parameters. The gradient describes the direction and rate of change of the loss function with the model parameters. The weight refers to the degree of influence of the model parameters on the loss function. It determines the speed at which the gradient descent algorithm updates the parameters. A larger gradient weight leads to faster parameter updates.
[0051] Grid search parameter tuning: It is a commonly used hyperparameter adjustment strategy. It is a process of finding the target hyperparameter combination by traversing all possible hyperparameter combinations. Specifically, we first need to define a hyperparameter space, which contains the possible value range of each hyperparameter. Then, we divide the space of these hyperparameters into a series of subspaces, each subspace corresponding to a set of hyperparameter combinations. Next, we apply each hyperparameter combination in each subspace to the model and measure the performance of the model. Finally, we will find the target hyperparameter combination as the basis for updating the model parameters.
[0052] Bayesian optimization: is a commonly used hyperparameter tuning strategy, a hyperparameter tuning strategy based on Bayes' theorem. Bayes' theorem is a theory of probability that is used to estimate the probability of an event, which can be used to infer the target optimization value of hyperparameters. In Bayesian parameter tuning, we first define a prior distribution that represents our knowledge of the hyperparameters. Then, we combine the prior distribution with the observed experimental results to obtain the posterior distribution. Through the posterior distribution, we can estimate the target optimization value of the hyperparameter.
[0053] Batch (mini-batch): In deep learning scenarios, by dividing the entire dataset into several small datasets, training the small dataset each time, it avoids the problem of huge computational complexity caused by all the data in the dataset participating in the training at one time. In addition, when performing gradient updates, the gradient direction does not differ significantly from the entire dataset, ensuring the training effect.
[0054] Graphics Processing Unit (GPU): A microprocessor used for graphics-related calculations. With the development of deep learning technology, it is widely used as computing hardware for model training because of its parallel structure, which can realize efficient matrix operations.
[0055] Tensor Processing Unit (TPU): A customized chip specifically for deep learning tasks that can provide higher efficiency and performance than GPU.
[0056] Field-Programmable Gate Array (FPGA): A programmable gate array that can be programmed to perform a variety of functions, including convolution operations in deep learning.
[0057] Application Specific Integrated Circuit (ASIC): An integrated circuit designed specifically for a specific application that can provide very high performance but at a higher cost.
[0058] The Central Processing Unit (CPU) can also be used for model training, but it is usually not as efficient as GPU and TPU.
[0059] Deep Self-Attention Model (Transformer Model): A deep learning architecture based on the attention mechanism (Attention) for processing sequence data such as natural language.
[0060] Bidirectional Encoder Representations from Transformers (BERT): A special Transformer model trained using a bidirectional Transformer encoder and large-scale unlabeled text data. BERT’s outstanding performance has made it a standard baseline for many natural language processing tasks.
[0061] Large Language Model (LLM): A deep learning model trained on a large corpus for natural language processing tasks. These models generally contain a multi-layer neural network, whose input is a text sequence for text generation, and the output is the task result text generated by performing a specific natural language processing task on the text sequence. Pre-training means that before a specific task, the model has been trained and learned to process a large amount of language data in advance. By pre-training the models, they can capture more complex language and semantic rules, thereby performing well in various natural language processing tasks and reducing the demand for large-scale data for specific tasks.
[0062] Distributed training: A method of training deep learning models using multiple distributed nodes, which can greatly increase training speed and reduce computation time.
[0063] Model Parallel (MP): A distributed training strategy that splits a model into multiple parts and deploys each part to different distributed nodes for training. This method can effectively solve the problem of too many parameters in the model.
[0064] Data Parallel (DP): A distributed training strategy that splits the entire sample set into multiple sample subsets and assigns each sample subset to a different distributed node for training. This technology can improve training efficiency and reduce memory usage.
[0065] Pipeline Parallel (PP): A distributed training strategy that is between model parallelism and data parallelism. The core idea is to decompose a large model into multiple layers and combine them into a pipeline to perform forward propagation and back propagation calculations, thereby reducing the memory usage of a single card and reducing communication overhead.
[0066] Remote Direct Memory Access (RDMA) is a technology used to directly read and write memory between remote computers, which can greatly improve the speed of data transmission in the network.
[0067] Network card: A hardware device that is mainly used to realize the physical connection between the computer and the network, and completes data transmission by sending and receiving data packets.
[0068] Peripheral Component Interconnect Express (PCIe): A physical connection structure between a group of nodes, usually consisting of a bus, which can be used to support data transmission between multiple devices.
[0069] Inter-GPU Express: A fast communication channel between two or more GPUs that provides higher bandwidth and lower latency for data transfer between GPUs.
[0070] Communication channel: A channel used to transmit data signals, which can be a physical channel or a logical channel. A physical channel is composed of a transmission medium and related communication equipment, and is used to transmit actual data signals; a logical channel refers to a logical path realized through an intermediate node on the basis of a physical channel, that is, a logical path formed between the sender and the receiver.
[0071] At present, the distributed training of deep learning models is mainly based on sample sets or model parameters of deep learning models to build multiple distributed data, and then distribute the multiple distributed data to different distributed nodes. Iterative training is performed on any distributed node. During any iterative training process, the training of the language model is completed in accordance with the method of propagation calculation and parameter update.
[0072] However, once a training anomaly occurs, such as hardware anomaly, system anomaly, network anomaly, or other unknown anomaly, distributed training needs to be re-executed without storing the updated model parameters, which is unbearable for deep learning models with large performance overhead. Although the updated model parameters can be stored at specific checkpoints (for example, after the parameter update of each iteration process), in the case of large deep learning model parameters, it is necessary to wait for the model parameters to be stored before continuing training. For deep learning models with model specifications reaching the migration level of tens of billions, such time overhead often reaches several minutes or even more than ten minutes, which determines that model parameters cannot be stored frequently. In this case, once a training anomaly occurs, the time overhead of resuming distributed training may reach several hours. Therefore, there is an urgent need for a deep learning model training method with high stability and high efficiency, which can achieve pre-storage of updated model parameters when training anomalies occur through low performance overhead, and no recalculation is required after resuming distributed training.
[0073] In response to the above problems, this specification provides a deep learning model training method. This specification also involves another deep learning model training method, a deep learning model training system, a deep learning model training device, another deep learning model training device, a computing device, a computer-readable storage medium and a computer program, which are described in detail one by one in the following embodiments.
[0074] See also Figure 1 , Figure 1 A flowchart of a deep learning model training method provided by an embodiment of this specification is shown, including the following specific steps:
[0075] Step 102: Obtain an initial deep learning model and sample data set.
[0076] The embodiments of this specification are applied to a server with a distributed training function, on which multiple distributed nodes and storage media are deployed, and any distributed node includes computing hardware for model training, such as a GPU, NPU, FPGA, ASIC, or CPU.
[0077] A deep learning model refers to a type of machine learning model based on a deep neural network structure, which is widely used in multiple fields such as vision, speech, and natural language processing. A deep learning model has large-scale model parameters, and includes but is not limited to: language processing models, image processing models, speech processing models, code processing models, etc. Taking the language processing model as an example, the language processing model can perform one or more natural language processing tasks, including but not limited to: machine translation tasks, speech recognition tasks, text analysis tasks, or text question and answer tasks. In terms of function, the language processing model can be regarded as: a translation model, a speech recognition model, a text analysis model, and a text question and answer model, etc., which are not limited here. In terms of model structure, the language processing model can be a Transformer model, a BERT model, a large language model, etc., which are not limited here.
[0078] The sample data set is a collection of sample data of the deep learning model used for training, and the sample data set includes large-scale sample data. The sample data set can be a labeled sample data set or an unlabeled sample data set. Depending on the training requirements, the sample data can be data of different modalities. For example, the deep learning model needs to be trained into a model with text processing function, sample text of text modality, and another example is that the deep learning model needs to be trained into a model with audio processing function, sample audio of audio modality, and another example is that the deep learning model needs to be trained into a model with image processing function, sample image of image modality, and another example is that the deep learning model needs to be trained into a model with numerical processing function, sample numerical value of numerical modality. The sample data set can be obtained from a sample database, such as an open source sample database, or it can be artificially constructed, such as generated using a generative model, and it can also be obtained from a historical database, such as obtaining historical query text and historical answer text from a historical database to construct a sample text set, which is not limited here.
[0079] For example, on a website of a large language model, a virtual character dialogue function needs to be added. By training the initial large language model, the trained target large language model has a text question-answering function and can perform text question-answering tasks. The initial large language model is obtained from the model library, and a sample text set is obtained from the open source sample database. The sample text set includes 10,000,000 sample question-answering text pairs.
[0080] Obtain the initial deep learning model and sample data set, which lays the foundation for the model and sample data for subsequent distributed training.
[0081] Step 104: According to the preset distributed training strategy, the deep learning model is distributedly trained based on the sample data set, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.
[0082] Distributed training is a training method that uses multiple distributed nodes to train deep learning models. Model training is iterative training, and one iteration process includes adjustment parameter calculation and model parameter update. Adjustment parameters are hyperparameters for adjusting model parameters, and adjustment parameter calculation is the calculation of hyperparameters for adjusting model parameters, including but not limited to: propagation calculation, probability calculation (updating model parameters through Bayesian optimization), grid calculation (updating model parameters through grid search) and diffusion derivation process (using diffusion model for forward diffusion process and reverse derivation process).
[0083] The distributed training strategy is a strategy that uses a distributed computing framework to split the large-scale training task of the deep learning model into multiple small-scale training tasks and distributes them to multiple distributed nodes for execution, including but not limited to: data parallel strategy, model parallel strategy and pipeline parallel strategy. The distributed training strategy is pre-set according to the model attributes, sample data sets and / or training tasks of the deep learning model. For example, the model layers of the deep learning model have execution restrictions in sequence, which makes it difficult to split the training, so the data parallel strategy is adopted. For another example, the data scale of the sample data set is too large, so the data parallel strategy is adopted. For another example, in the training task of large-scale text classification, the use of data parallel strategy may cause network bottlenecks, and the use of model parallel strategy may be too complex, so the pipeline parallel strategy is adopted. Different distributed training strategies correspond to different iterative characteristics of model training. For example, when the data parallel strategy is adopted, the amount of sample data on each distributed node is small, while the model parameter specifications are large, the time overhead of adjusting the parameter calculation at one time is high, and the number of times is small. When the model parallel strategy is adopted, the amount of sample data on each distributed node is large, while the model parameter specifications are small, the time overhead of adjusting the parameter calculation at one time is low, and the number of times is large. Different iteration characteristics determine different time costs for adjusting parameter calculations and storage times. It is necessary to finely determine the target storage parameters and overlap the model parameter storage process with the adjustment parameter calculation process in distributed training.
[0084] The model specification information is information about the model parameter specifications of the deep learning model, including but not limited to: model parameter quantity, benchmark model, etc. The model parameter quantity is the parameter specification quantity of the model parameter, for example, the model parameter quantity of the deep learning model is a specification of 10 billion. The benchmark model is a specific deep learning model architecture, for example, in natural language processing, the benchmark model of the deep learning model is the Transformer model, the BERT model, and the large language model, etc.
[0085] In the embodiments of this specification, the storage is the model parameter storage executed synchronously with the adjustment parameter calculation process, that is, the storage process and the adjustment parameter calculation process have a time overhead that is not much different. For example, the time overhead of the adjustment parameter calculation process is T, and the time overhead of the storage process is T'. If T is less than T', even during the model parameter update process, the storage is still being performed. Once a training anomaly occurs, it is difficult to resume model training. For details, see the following description. Taking propagation calculation as an example, the storage process can be in the forward propagation calculation process, in the reverse propagation calculation process, or in the forward propagation calculation process and the reverse propagation calculation process, which is not limited here.
[0086] The target storage parameters are configuration parameters for storing model parameters. The target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, including but not limited to: target storage model parameter specifications, target storage frequency, target storage read and write speed, target storage bandwidth, etc. After the distributed training strategy is determined and executed, the time cost of adjusting the parameter calculation process is difficult to change. In order to overlap the storage process of model parameters with the adjustment parameter calculation process in distributed training, it is necessary to control the time cost of storing model parameters through the target storage parameters.
[0087] It should be noted that if the model parameters are stored during the model parameter update process, the updated model parameters and the unupdated model parameters are mixed and stored. Once a training anomaly occurs, the model training cannot be restored by storing the mixed updated model parameters and the unupdated model parameters. For example, after completing the i-1th iteration training, the current model parameters are θ i-1 , perform the i-th iteration training, and in the process of adjusting the parameter calculation, adjust the current model parameter θ i-1 After the storage is completed, if a training anomaly occurs, the model parameter θ can be obtained i-1 Re-execute the i-th iteration training. If during the model parameter update process, the current model parameter θ i-1 For storage, some updated model parameters θ will be introduced i , and thus it is impossible to re-execute the i-th iteration training when a training anomaly occurs.
[0088] According to the preset distributed training strategy, the deep learning model is distributedly trained based on the sample data set. The specific method is: according to the preset distributed training strategy, multiple distributed data are constructed, the multiple distributed data are distributed to distributed nodes, and iterative training is performed, wherein the iterative training includes adjusting parameter calculation and model parameter update.
[0089] The model parameters of the deep learning model are stored according to the target storage parameters, specifically in the following manner: according to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium. The storage medium is a hardware device for storing model parameters, including but not limited to: computing hardware cache (GPU cache, NPU cache, FPGA cache, ASIC cache and CPU cache), memory, hard disk and distributed persistent storage array.
[0090] The target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy. The specific method is: based on the model specification information of the deep learning model and the preset distributed training strategy, time cost analysis is performed to obtain the target storage parameters. The time cost analysis is to analyze the time cost of adjusting the parameter calculation process and the model parameter storage process in iterative training, and then determine the target storage parameters that can be stored.
[0091] For example, the preset distributed training strategy is a data parallel strategy. The number of model parameters based on the large language model is 10 13 According to the level and data parallel strategy, the time cost analysis is performed to determine the time cost t1 of the forward propagation calculation process and the time cost t2 of the backward propagation calculation process of one iteration. Based on the time cost and the model parameter quantity, the target storage model parameter specifications are obtained as P1 and P2. According to the data parallel strategy, the 10,000,000 sample question-answer text pairs in the sample text set are divided. 64 distributed data are constructed and distributed to 64 distributed nodes. A GPU is deployed on each distributed node. On any distributed node, any distributed data is divided into 16 small batches, and the GPU is used to perform iterative training of 16 small batches. In the forward propagation calculation process and the backward propagation calculation process of each iterative training, according to the target storage model parameter specifications P1 and P2, the model parameters θ of the large language model are stored in the GPU cache, memory and hard disk of the distributed node in sequence. According to the above strategy, after the distributed training is completed, the target large language model with text question-answering function is obtained, and the target large language model is deployed on the cloud-side device of the website of the large language model to provide users with virtual character dialogue function.
[0092] In the embodiments of this specification, an initial deep learning model and a sample data set are obtained; distributed training is performed on the deep learning model based on the sample data set according to a preset distributed training strategy, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy. The target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, and the iterative law of distributed training is fully considered. In the process of calculating the adjustment parameters of distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, and the storage process of the model parameters and the adjustment parameter calculation process in the distributed training are overlapped, which fully saves performance overhead, and completes the real-time storage of the deep learning model parameters in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency.
[0093] In an optional embodiment of the present specification, in step 104, distributed training is performed on the deep learning model based on the sample data set according to a preset distributed training strategy, including the following specific steps:
[0094] According to a preset distributed training strategy, based on the deep learning model and the sample data set, multiple distributed data are constructed, wherein the preset distributed training strategy includes a model parallel training strategy or a data parallel training strategy;
[0095] Distribute multiple distributed data to each distributed node;
[0096] On a first distributed node, a propagation calculation is performed based on the distributed data to obtain a gradient weight, wherein the first distributed node is any one of the multiple distributed nodes;
[0097] Based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated, and when the preset training end conditions are met, a trained deep learning model is obtained.
[0098] Distributed data refers to partial data distributed to distributed nodes for execution. If the entire model training is understood as a large-scale training task, distributed data refers to the task data of the small-scale training tasks obtained by splitting the large-scale training task. Different distributed training strategies are used to construct different distributed data. For example, the model parallel strategy is used to split the model parameters of the deep learning model into multiple parts to construct distributed data. Another example is that the data parallel strategy is used to split the sample data set into multiple parts to construct distributed data.
[0099] Propagation calculation is the process of determining the gradient weight based on sample data using a deep learning model, including the forward propagation calculation process and the back propagation calculation process. Forward propagation calculation is the process of inputting sample data into a deep learning model and outputting predicted data. Back propagation calculation is the process of determining the loss value through predicted data, and then inputting the deep learning model back to determine the gradient weights of each model layer. Model parameter updating is the process of adjusting the model parameters of each model layer of the model based on the gradient weights. For example, in the forward propagation calculation process, sample data X is input into a large language model and predicted data Z is output. In the back propagation calculation process, the loss value is determined based on the predicted data Z, and the large language model is input back to determine the gradient weights of the n model layers of the large language model. for: During the model parameter update process, based on the gradient weight Adjust the model parameters of n model layers in a large language model.
[0100] Gradient weights are used to assign different gradient weights to the model parameter updates of each model layer during the reverse calculation propagation process, thereby controlling the update amplitude of the parameters of each model layer and the linear speed of the model parameter update. The larger the parameter update value of each model layer, the faster the parameter update, and vice versa.
[0101] The preset training end conditions are preset judgment conditions for stopping training, including but not limited to: a preset number of iterations, a preset loss value threshold, a preset training time, and a preset model convergence condition.
[0102] According to the preset distributed training strategy, multiple distributed data are constructed based on the deep learning model and the sample data set. The specific method is: according to the preset distributed training strategy, the model parameters and / or the sample data set of the deep learning model are divided to obtain multiple distributed data.
[0103] Propagation calculation is performed based on distributed data to obtain gradient weights. The specific method is as follows: sample data in the distributed data is input into the deep learning model, forward propagation calculation is performed to obtain predicted data, loss value is determined based on sample data and predicted data, loss value is reversely input into the deep learning model, back propagation calculation is performed to obtain gradient weights.
[0104] Based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated. Specifically, based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated by the gradient update method.
[0105] Exemplarily, according to the data parallel strategy, 10,000,000 sample question-answer text pairs in the sample text set are divided to obtain 64 distributed data. The 64 distributed data are distributed to 64 distributed nodes, each of which is deployed with a GPU. On any distributed node, any distributed data is divided into 16 small batches. During any iterative training process, the sample question text X in the small batch of distributed data is input into the large language model, and the forward propagation calculation is performed to obtain the predicted answer text Z'. Based on the sample answer text Z and the predicted answer text, the loss value Loss is determined, and the loss value is reversely input into the large language model, and the back propagation calculation is performed to obtain the gradient weight. Based on the gradient weights on each distributed node, the model parameters θ of the large language model are updated through the gradient update method. When all small batches of distributed data training are completed, the trained target large language model is obtained. The target large language model has a text question and answer function.
[0106] According to the preset distributed training strategy, multiple distributed data are constructed based on the deep learning model and the sample data set, wherein the preset distributed training strategy includes a model parallel training strategy or a data parallel training strategy; multiple distributed data are distributed to each distributed node; on the first distributed node, propagation calculation is performed based on the distributed data to obtain the gradient weight, wherein the first distributed node is any one of the multiple distributed nodes; based on the gradient weight on each distributed node, the model parameters of the deep learning model are updated, and when the preset training end conditions are met, a deep learning model that has been trained is obtained. According to the preset distributed training strategy, multiple distributed data are constructed, distributed training is performed on multiple distributed nodes, gradient weights are obtained, and the model parameters of the deep learning model are updated, thereby improving the efficiency of model training.
[0107] In an optional embodiment of the present specification, the distributed data includes multiple batches of distributed data;
[0108] On the first distributed node, a propagation calculation is performed based on the distributed data to obtain a gradient weight, including the following specific steps:
[0109] On the first distributed node, a propagation calculation is performed based on the distributed data of the current batch to obtain a gradient weight;
[0110] Correspondingly, after updating the model parameters of the deep learning model based on the gradient weights on each distributed node, the following specific steps are also included:
[0111] Update the distributed data of the current batch, return to the step of executing on the first distributed node, performing propagation calculation based on the distributed data of the current batch, and obtaining the gradient weight.
[0112] The current batch of distributed data is the distributed data of the batch for training the deep learning model in the current iterative training process. For example, on any distributed node, the distributed data is divided into 16 batches, and 16 iterative trainings need to be performed. One iterative training process includes a forward propagation calculation process, a backpropagation calculation process, and a parameter update process. Optionally, before the parameter update process, a communication process (a process of integrating gradient weights) is also included. The current is the i-th iterative training process, the current batch of distributed data is the i-th batch of distributed data, and the current model parameters are the model parameters θ after the i-1th update. i-1 Correspondingly, the distributed data of the current batch is updated, that is, the distributed data of the i-th batch is updated to the distributed data of the i+1-th batch.
[0113] Propagation calculation is performed based on the current batch of distributed data to obtain gradient weights. The specific method is: input the sample data in the current batch of distributed data into the deep learning model, perform forward propagation calculation, obtain predicted data, determine the loss value based on the sample data and the predicted data, reversely input the loss value into the deep learning model, perform back propagation calculation, and obtain gradient weights.
[0114] For example, the sample question text X in the distributed data of the i-th batch i Input large language model (the current model parameter is the model parameter θ after the i-1th update) i-1 ), perform forward propagation calculation, obtain prediction data prediction answer text Z' i , based on the sample answer text Z i And predict the answer text, determine the loss value Loss, reverse the loss value into the large language model, perform back propagation calculation, and obtain the gradient weight Based on the gradient weights on each distributed node, the model parameters of the large language model are updated from θ i-1 Update to θ i , update the current batch of distributed data from the i-th batch to the i+1-th batch, and continue to execute the sample question text X in the i+1-th batch of distributed data i+1 The steps of inputting a large language model are as follows: after completing the distributed data training of all small batches, a trained target large language model is obtained, and the target large language model has a text question-answering function.
[0115] On the first distributed node, the propagation calculation is performed based on the current batch of distributed data to obtain the gradient weight; the distributed data of the current batch is updated, and the step of performing the propagation calculation based on the current batch of distributed data on the first distributed node to obtain the gradient weight is returned. By dividing the distributed data into multiple batches and using multiple batches of distributed data to iteratively train the deep learning model, the problem of huge computational load caused by all distributed data participating in the distributed training at one time is avoided, which causes the problem of training bottleneck and ensures the training effect.
[0116] In an optional embodiment of the present specification, before performing propagation calculation on the distributed data of the current batch to obtain the gradient weight, the following specific steps are also included on the first distributed node:
[0117] For each batch of distributed data, based on the model specification information of the deep learning model and the preset distributed training strategy, the number and time overhead of propagation calculations are predicted;
[0118] Based on the number of propagation calculations and the time overhead, the target storage parameters corresponding to the propagation calculation process are determined.
[0119] In the embodiments of this specification, due to different distributed training strategies, the number of forward calculations and backward calculations required for each batch of distributed data and the corresponding time overhead are also different during the training of the deep learning model. For example, using a data parallel strategy, the amount of sample data on each distributed node is small, while the model parameter specifications are large, the time overhead of a propagation calculation is high, and the number of times is small, while using a model parallel strategy, the amount of sample data on each distributed node is large, while the model parameter specifications are small, the time overhead of a propagation calculation is low, and the number of times is large. It is necessary to finely determine the target storage parameters based on the number and time overhead of propagation calculations, and overlap the storage process of the model parameters in each iterative training process with the propagation calculation process in distributed training.
[0120] For each batch of distributed data, based on the model specification information of the deep learning model and the preset distributed training strategy, the number and time overhead of propagation calculations are predicted, which are specifically predicted by the prediction algorithm, for example, the torch.distributed module, the torch.profiler tool, and the tf.dataAPI interface.
[0121] Based on the number of propagation calculations and the time overhead, the target storage parameters corresponding to the propagation calculation process are determined. The specific method is: with the time overhead of the model parameter storage process not exceeding the time overhead of the propagation calculation process as the goal, based on the number of propagation calculations and the time overhead, the target storage parameters corresponding to the propagation calculation process are determined.
[0122] For example, through the torch.distributed module, for each batch of distributed data, the number of model parameters based on the large language model is 10 13 The level and data parallel strategy predicts the number of propagation calculations to be 16 times, the time overhead of the forward propagation calculation process is t1, and the time overhead of the backward propagation calculation process is t2. The time overhead T' of the model parameter storage process does not exceed the time overhead T=16*(t1+t2) of the propagation calculation process. The target storage model parameter specifications corresponding to the propagation calculation process are determined to be P1 and P2.
[0123] For each batch of distributed data, based on the model specification information of the deep learning model and the preset distributed training strategy, the number of propagation calculations and time overhead are predicted; based on the number of propagation calculations and time overhead, the target storage parameters corresponding to the propagation calculation process are determined. This ensures the feasibility of subsequent storage of model parameters and more accurately overlaps the storage process of model parameters with the propagation calculation process in distributed training.
[0124] In an optional embodiment of the present specification, before updating the model parameters of the deep learning model based on the gradient weights on each distributed node, the following specific steps are also included:
[0125] The gradient weights on each distributed node are integrated through the communication channels between the distributed nodes.
[0126] The communication channel between distributed nodes is a channel connection for data transmission between distributed nodes. It can be a physical connection, such as a network card, PCIe topology, or optical fiber, or a virtual connection, such as a high-speed channel between GPUs or between distributed nodes. The communication channel can be implemented through RDMA technology to improve the transmission speed.
[0127] For example, the high-speed channel constructed by the network cards on 64 distributed nodes uses RDMA technology to integrate the gradient weights on each distributed node. Based on the gradient weights on each distributed node, the model parameters θ of the large language model are updated through the gradient update method.
[0128] The communication channels between distributed nodes are used to integrate the gradient weights on each distributed node. The communication process of iterative training is completed in a centralized manner through the communication channels, which improves the efficiency and stability of model training.
[0129] In an optional embodiment of the present specification, any distributed node includes a first communication channel connected to a storage medium and a second communication channel connected to other distributed nodes;
[0130] In step 104, the model parameters of the deep learning model are stored according to the target storage parameters, including the following specific steps:
[0131] According to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the first communication channel;
[0132] Correspondingly, the gradient weights on each distributed node are integrated through the communication channel between the distributed nodes, including the following specific steps:
[0133] The gradient weights on each distributed node are integrated through the second communication channel.
[0134] In the embodiments of this specification, the storage medium is independent of the distributed nodes to avoid anomalies on any distributed node, causing data loss and ensuring the reliability of the entire distributed system. Therefore, the distributed nodes need to establish a communication channel with the storage medium for data storage. However, the resource performance on the distributed nodes is limited, and the communication channel is needed in the model parameter storage process. In the communication process, that is, to integrate the gradient weights on each distributed node, the communication channel is also needed. This is difficult to achieve for large-scale deep learning models. Therefore, it is necessary to distinguish the communication channels of the two processes and isolate the two data transmission processes to avoid introducing additional collective communications and ensure that no interference is caused to distributed training.
[0135] The first communication channel connecting the storage medium is a communication channel used for the model parameter storage process, including physical connections and virtual connections between distributed nodes and storage media, such as network cards, PCIe topology, optical fibers, read and write channels (data buses) between distributed nodes and storage media, etc.
[0136] The second communication channel connected to other distributed nodes is a communication channel used for the communication process, including physical connections and virtual connections between distributed nodes and storage media, such as network cards, PCIe topology, optical fibers, distributed nodes, high-speed channels between GPUs, high-speed channels between distributed nodes, etc. The second communication channel can be implemented through RDMA technology, which improves the transmission speed.
[0137] Exemplarily, the first communication channel is the physical connection between the network cards on the 64 distributed nodes and the distributed persistent storage array. According to the target storage model parameter specifications P1 and P2, the model parameters θ of the large language model are stored in the distributed persistent storage array through the first communication channel. The second communication channel is the physical connection between the network cards on the 64 distributed nodes, and the PCIe topology with GPUs inserted on each distributed node, with the high-speed channels between GPUs and the virtual connection between the high-speed channels between each distributed node. Through the second communication channel, using RDMA technology, the gradient weights on each distributed node are integrated Based on the gradient weights on each distributed node, the model parameters θ of the large language model are updated through the gradient update method.
[0138] According to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the first communication channel; the gradient weights on each distributed node are integrated through the second communication channel. The two data transmission processes of model parameter storage and communication are isolated to avoid introducing additional collective communications, ensuring that there is no interference with distributed training, improving the reliability of distributed training, and improving the effect of model training.
[0139] In an optional embodiment of the present specification, the propagation calculation includes forward propagation calculation and reverse propagation calculation;
[0140] In the propagation calculation process of the distributed training in step 104, before storing the model parameters of the deep learning model according to the target storage parameters, the following specific steps are also included:
[0141] Based on the model specification information of the deep learning model and the preset distributed training strategy, determine the first target storage parameter corresponding to the forward propagation calculation process and the second target storage parameter corresponding to the backward propagation calculation process;
[0142] Correspondingly, in step 104, during the propagation calculation process of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, including the following specific steps:
[0143] In the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameters;
[0144] During the back propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameters.
[0145] The propagation calculation process includes the forward propagation calculation process and the reverse propagation calculation process. The time overhead required for the two is different. Therefore, it is necessary to make a detailed distinction and determine the target storage parameters in the forward propagation calculation process and the target storage parameters in the reverse propagation calculation process.
[0146] The first target storage parameters are configuration parameters for storing model parameters during the forward propagation calculation process. The first target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, including but not limited to: first target storage model parameter specifications, first target storage frequency, first target storage read and write speed, first target storage bandwidth, etc.
[0147] The second target storage parameters are configuration parameters for storing model parameters during the back-propagation calculation process. The second target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, including but not limited to: second target storage model parameter specifications, second target storage frequency, second target storage read and write speed, second target storage bandwidth, etc.
[0148] Based on the model specification information of the deep learning model and the preset distributed training strategy, determine the first target storage parameter corresponding to the forward propagation calculation process and the second target storage parameter corresponding to the reverse propagation calculation process. The specific method is: for each batch of distributed data, based on the model specification information of the deep learning model and the preset distributed training strategy, predict the number and time overhead of forward propagation calculation and reverse propagation calculation, and based on the number and time overhead of forward propagation calculation and reverse propagation calculation, determine the first target storage parameter corresponding to the forward propagation calculation and the second target storage parameter corresponding to the reverse propagation calculation process.
[0149] During the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameters. The specific method is: during the forward propagation calculation process, the model parameters of the deep learning model are stored in the storage medium according to the first target storage parameters.
[0150] During the back propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameters. The specific method is: during the back propagation calculation process, the model parameters of the deep learning model are stored in the storage medium according to the second target storage parameters.
[0151] For example, through the multi-process parallel communication module, for each batch of distributed data, the model parameter amount based on the large language model is 10 13According to the level and data parallel strategy, the number of forward propagation calculations is predicted to be 16 times and the time cost is t1, the number of backward propagation calculations is predicted to be 16 times and the time cost is t2, and the time cost T' of the model parameter storage process does not exceed the time cost T = 16*(t1+t2) of the propagation calculation process. The first target storage model parameter specification corresponding to the forward propagation calculation process is determined to be P1 and the second target storage model parameter specification corresponding to the backward propagation calculation process is determined to be P2. During the forward propagation calculation process, according to the first target storage model parameter specification, the model parameters θ of the large language model are stored in the GPU cache, memory and distributed persistent storage array of the distributed nodes. During the backward propagation calculation process, according to the second target storage model parameter specification, the model parameters θ of the large language model are stored in the GPU cache, memory and distributed persistent storage array of the distributed nodes.
[0152] Based on the model specification information of the deep learning model and the preset distributed training strategy, the first target storage parameter corresponding to the forward propagation calculation process and the second target storage parameter corresponding to the reverse propagation calculation process are determined; in the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameter; in the reverse propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameter. The target storage parameters corresponding to forward propagation and reverse propagation are divided more finely, and the storage process of the model parameters is overlapped with the forward propagation calculation process and the reverse propagation calculation process in distributed training, so as to complete the real-time storage of the deep learning model parameters in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency.
[0153] In an optional embodiment of the present specification, storing the model parameters of the deep learning model according to the target storage parameters in step 104 includes the following specific steps:
[0154] Model parameters of the deep learning model are stored in multiple storage media according to target storage parameters and storage performance priorities of the multiple storage media.
[0155] Different storage media have different storage performance, including but not limited to: read and write speed and persistence. For example, the read and write speed of memory is higher than that of hard disk, but the persistence of memory is lower than that of hard disk. Compared with hard disk, memory has lower abnormal tolerance.
[0156] According to the target storage parameters and the storage performance priorities of multiple storage media, the model parameters of the deep learning model are stored in multiple storage media. The specific method is as follows: according to the target storage parameters and the storage performance priorities of multiple storage media, a storage strategy is determined, and according to the storage strategy, the model parameters of the deep learning model are stored in multiple storage media. The embodiment of this specification provides a storage strategy: establish a storage medium hierarchy: from top to bottom, they are GPU cache, memory and hard disk, 1. Use higher-level storage media with faster read and write speeds as much as possible to maximize the parameter specification and abnormal recovery capabilities of model parameters; 2. Even if the upper-level storage medium is unavailable, the lower-level, more persistent storage medium can still be relied on to ensure the persistent storage of the current model parameters; 3. Save overhead through asynchronous execution between upper and lower layers.
[0157] Optionally, according to the target storage parameters, the model parameters of the deep learning model are stored in a multi-level storage medium, including the following specific steps: according to the target storage parameters, the model parameters of the deep learning model are stored in a first storage medium and a second storage medium respectively, wherein the reading and writing speed of the first storage medium is faster than that of the second storage medium.
[0158] Exemplarily, according to the target storage model parameter specifications P1 and P2, the model parameters θ of the large language model are stored in the GPU cache, memory and hard disk of the distributed nodes respectively.
[0159] Optionally, according to the target storage parameters, the model parameters of the deep learning model are stored in a multi-level storage medium, including the following specific steps: according to the target storage parameters, the model parameters of the deep learning model are stored in a first storage medium, so that the model parameters are transferred from the first storage medium to the second storage medium, wherein the reading and writing speed of the first storage medium is faster than that of the second storage medium.
[0160] Exemplarily, according to the target storage model parameter specifications P1 and P2, the model parameters θ of the large language model are stored in the GPU cache, so that the model parameters are transferred from the GPU cache to the memory, so that the model parameters are transferred from the memory to the hard disk.
[0161] According to the target storage parameters and the storage performance priorities of multiple storage media, the model parameters of the deep learning model are stored in multiple storage media. The storage performance of different storage media is fully utilized to achieve high-frequency model parameter storage, and the current model parameters are stored, saving overhead.
[0162] In an optional embodiment of the present specification, the method further includes the following specific steps:
[0163] Upon receiving a request to resume training, obtaining stored target model parameters, wherein the request to resume training is generated after determining that the training anomaly of the distributed training has been recovered, and the target model parameters are model parameters stored before the training anomaly occurs;
[0164] Resume distributed training of deep learning models based on target model parameters.
[0165] Training exceptions are abnormal situations that occur during model training. Abnormal situations may cause model training to fail to run normally or affect the performance and accuracy of the trained model, including but not limited to: hardware abnormalities, system abnormalities, network abnormalities, or other unknown abnormalities. Training abnormality recovery refers to taking corresponding measures to resume model training when training abnormalities occur. Resume training request is an instruction request for resuming model training.
[0166] The target model parameters are the model parameters stored before the training anomaly occurs. For example, after completing the i-1th iteration training, the target model parameters are θ i-1 , perform the i-th iteration training, during the propagation calculation process, the target model parameter θ i-1 After the storage is completed, if a training anomaly occurs, the model parameter θ can be obtained i-1 Re-execute the i-th iteration training.
[0167] The stored target model parameters are obtained by: obtaining the stored target model parameters from a storage medium.
[0168] Exemplarily, when a request to resume training is received, the stored target model parameters θ are obtained from the hard disk. Based on the target model parameters θ, the distributed training of the large language model is resumed to obtain a trained target large language model, which has a text question-answering function.
[0169] When a request to resume training is received, the stored target model parameters are obtained, where the request to resume training is generated after determining that the training anomaly of distributed training has been recovered, and the target model parameters are the model parameters stored before the training anomaly occurs; based on the target model parameters, the distributed training of the deep learning model is resumed. This increases the fault tolerance of deep learning model training, avoids retraining of deep learning models, ensures stability while ensuring training efficiency, and reduces training costs.
[0170] Figure 2 A system architecture diagram of a deep learning model training method provided by an embodiment of this specification is shown. Figure 2 As shown:
[0171] The system architecture includes multi-layer storage media, from top to bottom: GPU cache, memory, and hard disk. The multi-layer storage media has higher persistence from top to bottom, and the read and write speeds are faster from bottom to top. During distributed training, the model parameters of the deep learning model are stored from the GPU to the GPU cache for non-persistent high-speed reading and writing. The model parameters of the deep learning model are stored from the GPU to the memory for non-persistent high-speed reading and writing. The model parameters of the deep learning model are transferred from the memory to the hard disk for persistent storage. When training needs to be resumed, the model parameters of the deep learning model are obtained from the multi-layer storage media and the model training is performed again.
[0172] Figure 3 A flowchart of a deep learning model training method provided by an embodiment of this specification is shown. Figure 3 As shown:
[0173] At present, in the iterative training of the model, each iteration includes propagation calculation, parameter update and parameter storage. In the i-1th iteration, the propagation calculation and parameter update are completed, and the updated model parameters are stored. The i-th iteration is started, and the propagation calculation and parameter update are also completed, and the updated model parameters are stored.
[0174] This manual Figure 1 In the embodiment, each iteration includes propagation calculation, communication integration and parameter updating process. In the propagation calculation process of the i-1th iteration, the model parameters are parameter stored, and then the communication integration and parameter updating process are performed, and the i-th iteration is started. In the propagation calculation process of the i-th iteration, the model parameters are parameter stored, and then the communication integration and parameter updating process are performed.
[0175] Compared with the two, it saves time and improves the efficiency of model training.
[0176] Figure 4 A schematic diagram of a deep learning model training method provided by an embodiment of this specification is shown, such as Figure 4 As shown:
[0177] The system includes multiple distributed nodes (two distributed nodes in the figure) and storage media (distributed persistent storage array in the figure), and any distributed node includes memory, multiple graphics processing units, a first network card, and a second network card. A high-speed communication channel is established between each graphics processing unit in the distributed node, a first communication channel is established between the first network card on the distributed node and the distributed storage node, and a second communication channel is established between the second network cards of each distributed node.
[0178] See also Figure 5 , Figure 5 A flowchart of another deep learning model training method provided by an embodiment of the present specification is shown, which is applied to a cloud-side device, where the cloud-side device includes multiple distributed nodes and a storage medium; the method includes the following specific steps:
[0179] Step 502: Obtain an initial deep learning model and sample data set.
[0180] Step 504: According to the preset distributed training strategy, multiple distributed nodes are called to perform distributed training on the deep learning model based on the sample data set, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored in the storage medium according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.
[0181] Step 506: When an abnormality in the deep learning model training is identified, trigger multiple distributed nodes to stop the distributed training of the deep learning model, and determine the target model parameters currently stored in the storage medium.
[0182] Step 508: When a request to resume training is received, the target model parameters are obtained from the storage medium.
[0183] Step 510: Call multiple distributed nodes to perform distributed training on the deep learning model based on target model parameter recovery.
[0184] The embodiments of this specification are applied to a cloud-side device with a distributed training function. The cloud-side device is a network cloud device, a virtual device, and is composed of multiple distributed nodes and storage media. Any distributed node includes computing hardware for model training, such as GPU, NPU, FPGA, ASIC or CPU.
[0185] The embodiments of this specification are similar to the above Figure 1 The embodiments of the specification are based on the same inventive concept, and the specific methods of steps 502, 504, 508 and 510 are described in the above Figure 1 The detailed description is given in the embodiments of the specification and will not be repeated here.
[0186] For example, on a website of a large language model, it is currently necessary to add a virtual character dialogue function. By training the initial large language model, the trained target large language model has a text question-answering function and can perform text question-answering tasks. The initial large language model is obtained from the model library, and the sample text set is obtained from the open source sample database. The sample text set includes 10,000,000 sample question-answer text pairs. The preset distributed training strategy is the data parallel strategy. The model parameter quantity based on the large language model is 10 13Level and data parallel strategy, perform time cost analysis, determine the time cost of one iteration of the forward propagation calculation process is t1 and the time cost of the back propagation calculation process is t2, based on the time cost and model parameter quantity, obtain the target storage model parameter specifications as P1 and P2. According to the data parallel strategy, the 10,000,000 sample question-answer text pairs in the sample text set are divided. Construct 64 distributed data, distribute the 64 distributed data to 64 distributed nodes, each distributed node is deployed with a GPU, and on any distributed node, divide any distributed data into 16 small batches, and use the GPU to perform iterative training of 16 small batches. In the forward propagation calculation process and the back propagation calculation process of each iterative training, the model parameters θ of the large language model are stored in the GPU cache, memory and hard disk of the distributed nodes in turn according to the target storage model parameter specifications P1 and P2. When the large language model training abnormality is identified, the 64 distributed nodes are triggered to stop the distributed training of the large language model, and the target model parameters θ currently stored in the hard disk are determined. When a request to resume training is received, the stored target model parameters θ are obtained from the hard disk. Based on the target model parameters θ, the distributed training of the large language model is resumed to obtain the trained target large language model. The target large language model has a text question-and-answer function. The target large language model is deployed on the cloud-side device of the website of the large language model to provide users with a virtual character dialogue function.
[0187] In an embodiment of the present specification, an initial deep learning model and a sample data set are obtained; according to a preset distributed training strategy, multiple distributed nodes are called to perform distributed training on the deep learning model based on the sample data set, and during the calculation of adjustment parameters for the distributed training, the model parameters of the deep learning model are stored in a storage medium according to target storage parameters, wherein the target storage parameters are determined based on model specification information of the deep learning model and a preset distributed training strategy; when an abnormality in the deep learning model training is identified, multiple distributed nodes are triggered to stop the distributed training of the deep learning model, and the target model parameters currently stored in the storage medium are determined; when a request to resume training is received, the target model parameters are obtained from the storage medium; and multiple distributed nodes are called to resume distributed training of the deep learning model based on the target model parameters. The target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, and the iterative law of distributed training is fully considered. In the process of calculating the adjustment parameters of distributed training, the model parameters of the deep learning model are stored in the storage medium according to the target storage parameters, so that the storage process of the model parameters and the adjustment parameter calculation process in distributed training are overlapped, which fully saves the performance overhead. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency. When the deep learning model training anomaly is identified, the stored target model parameters are obtained from the storage medium, and the distributed training of the deep learning model is restored, which increases the fault tolerance of the deep learning model training and avoids re-training of the deep learning model. It has stability while ensuring training efficiency and reducing training costs.
[0188] In an optional embodiment of the present specification, the storage medium includes a plurality of storage media with different storage performances;
[0189] In step 504, during the distributed training adjustment parameter calculation process, the model parameters of the deep learning model are stored in the storage medium according to the target storage parameters, including the following specific steps:
[0190] storing model parameters of the deep learning model in multiple storage media according to target storage parameters and storage performance priorities of the multiple storage media;
[0191] Correspondingly, obtaining the target model parameters from the storage medium in step 508 includes the following specific steps:
[0192] Acquiring target model parameters from a first storage medium;
[0193] If not obtained, the target model parameters are obtained from the second storage medium, wherein the storage performance priority of the first storage medium is higher than that of the second storage medium.
[0194] The steps of storing the model parameters of the deep learning model to multiple storage media according to the target storage parameters and the storage performance priorities of multiple storage media have been described above. Figure 1 The detailed description is given in the embodiments of the specification and will not be repeated here.
[0195] Considering that the storage performance priority of the first storage medium is higher than that of the second storage medium, Figure 2 For example, the first storage medium is memory, and the second storage medium is a hard disk. The read and write speed of the memory is higher than that of the hard disk. If it can be obtained, the training efficiency is higher than that of the hard disk. However, the memory is non-persistent storage. Therefore, the target model parameters may not be obtained and need to be obtained from the hard disk.
[0196] Exemplarily, the stored target model parameter θ is obtained from the memory, and if not obtained, the target model parameter θ is obtained from the hard disk.
[0197] The target model parameters are obtained from the first storage medium; if the target model parameters are not obtained, the target model parameters are obtained from the second storage medium, wherein the storage performance priority of the first storage medium is higher than that of the second storage medium. The storage performance priority differences of the storage mediums are fully utilized. In the event of training anomalies, the target model parameters are obtained from the storage medium with high storage performance priority first, and the storage medium with high storage performance priority is ensured to have persistent storage of the target model parameters, which improves the efficiency of model training while ensuring the reliability of model training.
[0198] In an optional embodiment of the present specification, the cloud-side device further includes a first communication channel connected to each distributed node and a second communication channel connected to the storage medium;
[0199] In step 504, according to the preset distributed training strategy, multiple distributed nodes are called to perform distributed training on the deep learning model based on the sample data set, including the following specific steps:
[0200] According to a preset distributed training strategy, multiple distributed nodes are called through the first communication channel to perform distributed training on the deep learning model based on the sample data set;
[0201] Correspondingly, in step 504, storing the model parameters of the deep learning model to the storage medium according to the target storage parameters includes the following specific steps:
[0202] According to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the second communication channel.
[0203] The second communication channel connected to the storage medium is a communication channel used for the model parameter storage process, including physical channels and logical channels between distributed nodes and storage media, such as network cards, PCIe topology, optical fibers, read and write channels (data buses) between distributed nodes and storage media, etc.
[0204] The first communication channel connected to other distributed nodes is a communication channel used for the communication process, including a physical channel and a logical channel between a distributed node and a storage medium, for example, a network card, a PCIe topology, an optical fiber, a distributed node, a high-speed channel between GPUs, a high-speed channel between each distributed node, etc. The first communication channel can be implemented by RDMA technology, which improves the transmission speed.
[0205] Exemplarily, the second communication channel is a physical channel between the network cards on the 64 distributed nodes and the distributed persistent storage array. According to the target storage model parameter specification P, the model parameter θ of the large language model is stored in the distributed persistent storage array through the second communication channel. The first communication channel is a physical channel between the network cards on the 64 distributed nodes, and a PCIe topology with a GPU inserted on each distributed node, with a logical channel of a high-speed channel between GPUs and a high-speed channel between each distributed node. Through the first communication channel, the gradient weights on each distributed node are integrated using RDMA technology. Based on the gradient weights on each distributed node, the model parameters θ of the large language model are updated through the gradient update method.
[0206] According to the preset distributed training strategy, multiple distributed nodes are called through the first communication channel to perform distributed training on the deep learning model based on the sample data set; according to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the second communication channel. The two data transmission processes of model parameter storage and communication are isolated to avoid the introduction of additional collective communications, ensure that there is no interference with distributed training, improve the reliability of distributed training, and improve the effect of model training.
[0207] Corresponding to the above method embodiment, this specification also provides a deep learning model training system embodiment, Figure 6 FIG. 1 shows a schematic diagram of a deep learning model training system provided by an embodiment of this specification. Figure 6 As shown, the system includes a management and control unit 602 and multiple distributed nodes, the multiple distributed nodes include a first distributed node 604, and the first distributed node 604 is any one of the multiple distributed nodes;
[0208] The management and control unit 602 is used to obtain an initial deep learning model and a sample data set, build multiple distributed data based on the deep learning model and the sample data set according to a preset distributed training strategy, and distribute the multiple distributed data to each distributed node;
[0209] The first distributed node 604 is used to perform distributed training on the deep learning model based on the sample data set; and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored therein according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.
[0210] Optionally, the system further includes a persistent storage medium, and the first distributed node 604 includes a non-persistent storage medium;
[0211] Correspondingly, the first distributed node 604 is also used to store the model parameters of the deep learning model to a non-persistent storage medium according to the target storage parameters during the adjustment parameter calculation process of distributed training, so as to transfer the model parameters of the deep learning model from the non-persistent storage medium to the persistent storage medium.
[0212] Optionally, the adjustment parameter is a gradient weight, and the adjustment parameter calculation is a propagation calculation; the first distributed node 604 further includes a first communication channel connected to each distributed node and a second communication channel connected to the storage medium;
[0213] Correspondingly, the first distributed node 604 is further used to integrate the gradient weights on each distributed node through the first communication channel;
[0214] Correspondingly, the first distributed node 604 is also used to send the model parameters of the deep learning model from the non-persistent storage medium to the persistent storage medium for storage through the second communication channel.
[0215] In the embodiments of this specification, the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, and the iterative law of distributed training is fully considered. In the adjustment parameter calculation process of distributed training, the model parameters of the deep learning model are stored according to the target storage parameters. The storage process of the model parameters and the adjustment parameter calculation process in distributed training are overlapped, which fully saves performance overhead. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency.
[0216] The above is a schematic scheme of a deep learning model training system of this embodiment. It should be noted that the technical scheme of the deep learning model training system and the technical scheme of the deep learning model training method described above belong to the same concept, and the details not described in detail in the technical scheme of the deep learning model training system can be found in the description of the technical scheme of the deep learning model training method described above.
[0217] Corresponding to the above method embodiment, this specification also provides a deep learning model training device embodiment, Figure 7 FIG. 1 is a schematic diagram showing a structure of a deep learning model training device provided by an embodiment of the present specification. Figure 7 As shown, the device comprises:
[0218] A first acquisition module 702 is configured to acquire an initial deep learning model and a sample data set;
[0219] The first training module 704 is configured to perform distributed training on the deep learning model based on the sample data set according to a preset distributed training strategy, and store the model parameters of the deep learning model according to the target storage parameters during the calculation of the adjustment parameters of the distributed training, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.
[0220] Optionally, the first training module 704 is further configured to:
[0221] According to the preset distributed training strategy, multiple distributed data are constructed based on the deep learning model and the sample data set, wherein the preset distributed training strategy includes a model parallel training strategy or a data parallel training strategy; the multiple distributed data are distributed to each distributed node; on the first distributed node, propagation calculation is performed based on the distributed data to obtain the gradient weight, wherein the first distributed node is any one of the multiple distributed nodes; based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated, and when the preset training end conditions are met, a trained deep learning model is obtained.
[0222] Optionally, the distributed data includes multiple batches of distributed data;
[0223] The first training module 704 is further configured to:
[0224] On the first distributed node, a propagation calculation is performed based on the distributed data of the current batch to obtain a gradient weight;
[0225] Correspondingly, the device also includes:
[0226] The iteration module is configured to update the distributed data of the current batch, return to the step of performing propagation calculation on the first distributed node based on the distributed data of the current batch, and obtain the gradient weight.
[0227] Optionally, the device further comprises:
[0228] The storage parameter determination module is configured to predict the number of propagation calculations and the time overhead for each batch of distributed data based on the model specification information of the deep learning model and the preset distributed training strategy; based on the number of propagation calculations and the time overhead, determine the target storage parameters corresponding to the propagation calculation process.
[0229] Optionally, the device further comprises:
[0230] The gradient weight integration module is configured to integrate the gradient weights on each distributed node through the communication channel between the distributed nodes.
[0231] Optionally, any distributed node includes a first communication channel connected to a storage medium and a second communication channel connected to other distributed nodes;
[0232] The first training module 704 is further configured to:
[0233] According to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the first communication channel;
[0234] Correspondingly, the gradient weight integration module is further configured as:
[0235] The gradient weights on each distributed node are integrated through the second communication channel.
[0236] Optionally, the propagation calculation includes forward propagation calculation and backward propagation calculation;
[0237] The device also includes:
[0238] A forward and reverse storage parameter determination module is configured to determine a first target storage parameter corresponding to a forward propagation calculation process and a second target storage parameter corresponding to a reverse propagation calculation process based on model specification information of a deep learning model and a preset distributed training strategy;
[0239] Correspondingly, the first training module 704 is further configured as follows:
[0240] During the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameters; during the backward propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameters.
[0241] Optionally, the device further comprises:
[0242] The training resumption module is configured to obtain the stored target model parameters upon receiving a training resumption request, wherein the training resumption request is generated after determining that the distributed training has recovered from a training anomaly, and the target model parameters are the model parameters stored before the training anomaly occurs; based on the target model parameters, the distributed training of the deep learning model is resumed.
[0243] In the embodiments of this specification, the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, and the iterative law of distributed training is fully considered. In the propagation calculation process of distributed training, the model parameters of the deep learning model are stored according to the target storage parameters. The storage process of the model parameters is overlapped with the propagation calculation process in distributed training, which fully saves performance overhead. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency.
[0244] The above is a schematic scheme of a deep learning model training device of this embodiment. It should be noted that the technical scheme of the deep learning model training device and the technical scheme of the deep learning model training method described above belong to the same concept, and the details not described in detail in the technical scheme of the deep learning model training device can be found in the description of the technical scheme of the deep learning model training method described above.
[0245] Corresponding to the above method embodiment, this specification also provides a deep learning model training device embodiment, Figure 8 FIG. 2 shows a schematic diagram of the structure of another deep learning model training device provided by an embodiment of the present specification. Figure 8 As shown, it is applied to a cloud side device, and the cloud side device includes multiple distributed nodes and a storage medium; the device includes:
[0246] A second acquisition module 802 is configured to acquire an initial deep learning model and a sample data set;
[0247] The second training module 804 is configured to call multiple distributed nodes according to a preset distributed training strategy, perform distributed training on the deep learning model based on the sample data set, and store the model parameters of the deep learning model to the storage medium according to the target storage parameter during the adjustment parameter calculation process of the distributed training, wherein the target storage parameter is determined based on the model specification information of the deep learning model and the preset distributed training strategy;
[0248] The stop module 806 is configured to trigger the multiple distributed nodes to stop the distributed training of the deep learning model when an abnormality in the deep learning model training is identified, and determine the target model parameters currently stored in the storage medium;
[0249] The parameter acquisition module 808 is configured to acquire the target model parameters from the storage medium when receiving the request to resume training;
[0250] The recovery module 810 is configured to call multiple distributed nodes to perform distributed training on the deep learning model based on target model parameter recovery.
[0251] Optionally, the storage medium includes a plurality of storage media with different storage performances;
[0252] The second training module 804 is further configured to:
[0253] storing model parameters of the deep learning model in multiple storage media according to target storage parameters and storage performance priorities of the multiple storage media;
[0254] Correspondingly, the parameter acquisition module 808 is further configured as follows:
[0255] The target model parameters are obtained from the first storage medium; if the target model parameters are not obtained, the target model parameters are obtained from the second storage medium, wherein the storage performance priority of the first storage medium is higher than that of the second storage medium.
[0256] Optionally, the cloud-side device further includes a first communication channel connected to each distributed node and a second communication channel connected to the storage medium;
[0257] The second training module 804 is further configured to:
[0258] According to the preset distributed training strategy, multiple distributed nodes are called through the first communication channel to perform distributed training on the deep learning model based on the sample data set; according to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the second communication channel.
[0259] In the embodiments of this specification, target storage parameters are determined based on model specification information of the deep learning model and a preset distributed training strategy, and the iterative law of distributed training is fully considered. In the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored in a storage medium according to the target storage parameters, so that the storage process of the model parameters and the adjustment parameter calculation process in the distributed training are overlapped, and the performance overhead is fully saved. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency. When the deep learning model training anomaly is identified, the stored target model parameters are obtained from the storage medium, and the distributed training of the deep learning model is restored, which increases the fault tolerance of the deep learning model training and avoids re-training of the deep learning model. While having stability, it ensures training efficiency and reduces training costs.
[0260] The above is a schematic scheme of a deep learning model training device of this embodiment. It should be noted that the technical scheme of the deep learning model training device and the technical scheme of the deep learning model training method described above belong to the same concept, and the details not described in detail in the technical scheme of the deep learning model training device can be found in the description of the technical scheme of the deep learning model training method described above.
[0261] Fig. 9 The structure block diagram of a computing device provided by an embodiment of the present specification is shown. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and the database 950 is used to store data.
[0262] The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, and a Near Field Communication (NFC).
[0263] In one embodiment of the present specification, the above components of the computing device 900 and Fig. 9 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig. 9 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0264] The computing device 900 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 900 may also be a mobile or stationary server.
[0265] Among them, the processor 920 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned deep learning model training method.
[0266] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned deep learning model training method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above-mentioned deep learning model training method.
[0267] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned deep learning model training method.
[0268] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned deep learning model training method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above-mentioned deep learning model training method.
[0269] An embodiment of the present specification also provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned deep learning model training method.
[0270] The above is a schematic scheme of a computer program of this embodiment. It should be noted that the technical scheme of the computer program and the technical scheme of the above-mentioned deep learning model training method belong to the same concept, and the details not described in detail in the technical scheme of the computer program can be found in the description of the technical scheme of the above-mentioned deep learning model training method.
[0271] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0272] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0273] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0274] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0275] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A deep learning model training method, comprising: Obtain an initial deep learning model and sample dataset; According to the preset distributed training strategy, the deep learning model is distributedly trained based on the sample data set, and in the adjustment parameter calculation process of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, the adjustment parameter calculation is a propagation calculation, and the target storage parameters are determined based on the number and time overhead of the propagation calculation, with the goal that the time overhead of the model parameter storage process does not exceed the time overhead of the propagation calculation process.
2. The method according to claim 1, wherein the adjustment parameter is a gradient weight; The performing distributed training on the deep learning model based on the sample data set according to a preset distributed training strategy includes: According to a preset distributed training strategy, based on the deep learning model and the sample data set, a plurality of distributed data are constructed, wherein the preset distributed training strategy includes a model parallel training strategy or a data parallel training strategy; Distributing the plurality of distributed data to each distributed node; On a first distributed node, performing propagation calculation based on distributed data to obtain a gradient weight, wherein the first distributed node is any one of the multiple distributed nodes; Based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated, and when the preset training end conditions are met, the trained deep learning model is obtained.
3. The method according to claim 2, wherein the distributed data comprises a plurality of batches of distributed data; The step of performing propagation calculation on the first distributed node based on the distributed data to obtain the gradient weight includes: On the first distributed node, a propagation calculation is performed based on the distributed data of the current batch to obtain a gradient weight; After updating the model parameters of the deep learning model based on the gradient weights on each distributed node, the method further includes: The distributed data of the current batch is updated, and the step of performing calculation based on the distributed data of the current batch on the first distributed node to obtain the gradient weight is returned.
4. The method according to claim 3, before performing calculation on the distributed data of the current batch on the first distributed node to obtain the gradient weight, further comprising: For each batch of distributed data, based on the model specification information of the deep learning model and the preset distributed training strategy, predict the number of calculations and time overhead; Based on the number of times and time cost of the calculation, a target storage parameter corresponding to the calculation process is determined.
5. The method according to claim 2, before updating the model parameters of the deep learning model based on the gradient weights on each distributed node, further comprising: The gradient weights on the distributed nodes are integrated through the communication channels between the distributed nodes.
6. The method according to claim 5, any distributed node comprises a first communication channel connected to the storage medium and a second communication channel connected to other distributed nodes; The storing of the model parameters of the deep learning model according to the target storage parameters includes: According to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the first communication channel; The step of integrating the gradient weights on the distributed nodes through the communication channels between the distributed nodes includes: The gradient weights on the distributed nodes are integrated through the second communication channel.
7. The method according to claim 1, wherein the propagation calculation comprises a forward propagation calculation and a backward propagation calculation; In the calculation process of the distributed training, before storing the model parameters of the deep learning model according to the target storage parameters, the method further includes: Based on the model specification information of the deep learning model and the preset distributed training strategy, determine a first target storage parameter corresponding to a forward propagation calculation process and a second target storage parameter corresponding to a backward propagation calculation process; The storing of the model parameters of the deep learning model according to the target storage parameters during the calculation process of the distributed training includes: In the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameters; During the back propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameters.
8. The method according to claim 1, further comprising: Upon receiving a request to resume training, obtaining stored target model parameters, wherein the request to resume training is generated after determining that a training anomaly of distributed training has been recovered, and the target model parameters are model parameters stored before the training anomaly occurs; Based on the target model parameters, resume distributed training of the deep learning model.
9. A deep learning model training method, applied to a cloud-side device, wherein the cloud-side device includes a plurality of distributed nodes and a storage medium; the method includes: Obtain an initial deep learning model and sample dataset; According to the preset distributed training strategy, the multiple distributed nodes are called to perform distributed training on the deep learning model based on the sample data set, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored in the storage medium according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, the adjustment parameter calculation is a propagation calculation, and the target storage parameters are determined based on the number of times and time overhead of the propagation calculation with the goal that the time overhead of the model parameter storage process does not exceed the time overhead of the propagation calculation process; In the case of identifying an abnormality in the deep learning model training, triggering the multiple distributed nodes to stop the distributed training of the deep learning model, and determining the target model parameters currently stored in the storage medium; Upon receiving a request to resume training, obtaining the target model parameters from the storage medium; The multiple distributed nodes are called to perform distributed training on the deep learning model based on the target model parameter recovery.
10. The method according to claim 9, wherein the storage medium comprises a plurality of storage media with different storage performances; The step of storing the model parameters of the deep learning model in the storage medium according to the target storage parameters during the calculation process of the distributed training includes: storing the model parameters of the deep learning model in the multiple storage media according to the target storage parameters and the storage performance priorities of the multiple storage media; The acquiring the target model parameters from the storage medium includes: Acquiring the target model parameters from a first storage medium; If not obtained, the target model parameters are obtained from a second storage medium, wherein the storage performance priority of the first storage medium is higher than that of the second storage medium.
11. The method according to claim 9, wherein the cloud-side device further comprises a first communication channel connected to each distributed node and a second communication channel connected to the storage medium; According to a preset distributed training strategy, calling the multiple distributed nodes, and performing distributed training on the deep learning model based on the sample data set, including: According to a preset distributed training strategy, calling the plurality of distributed nodes through the first communication channel, and performing distributed training on the deep learning model based on the sample data set; The storing the model parameters of the deep learning model to the storage medium according to the target storage parameters includes: According to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the second communication channel.
12. A deep learning model training system, the system comprising a management and control unit and a plurality of distributed nodes, the plurality of distributed nodes comprising a first distributed node, the first distributed node being any one of the plurality of distributed nodes; The control unit is used to obtain an initial deep learning model and a sample data set, construct a plurality of distributed data based on the deep learning model and the sample data set according to a preset distributed training strategy, and distribute the plurality of distributed data to each distributed node; The first distributed node is used to perform distributed training on the deep learning model based on the sample data set; And in the adjustment parameter calculation process of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, the adjustment parameter calculation is a propagation calculation, and the target storage parameters are determined based on the number and time overhead of the propagation calculation, with the goal that the time overhead of the model parameter storage process does not exceed the time overhead of the propagation calculation process.
13. The system of claim 12, wherein the system further comprises a persistent storage medium, and the first distributed node comprises a non-persistent storage medium; The first distributed node is also used to store the model parameters of the deep learning model to the non-persistent storage medium according to the target storage parameters during the adjustment parameter calculation process of the distributed training, so as to transfer the model parameters of the deep learning model from the non-persistent storage medium to the persistent storage medium.
14. The system according to claim 13, wherein the adjustment parameter is a gradient weight; The first distributed node also includes a first communication channel connected to each distributed node and a second communication channel connected to the storage medium; The first distributed node is further used to integrate the gradient weights on each distributed node through the first communication channel; The first distributed node is also used to send the model parameters of the deep learning model from the non-persistent storage medium to the persistent storage medium for storage through the second communication channel.
15. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 11 are implemented.
16. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for neural network machine learning model training
CN109754060A
Deep learning model training method and system
CN111788585A
Cited By
Deep learning model training method and deep learning model training system
EP4818983A1
Deep learning model training method and deep learning model training system
WO2025112801A1