Deep learning model training method and deep learning model training system

In the distributed training of deep learning models, the model parameters are stored in real time based on the model specification information and the target storage parameters determined by the distributed training strategy, which solves the progress loss and performance overhead problems caused by training exceptions, and realizes an efficient and fault-tolerant training process.

WO2025112801A1PCT designated stage expired Publication Date: 2025-06-05HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD

Patent Information

Application Number
PCT/CN2024/118478
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-09-12
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

During distributed training, training exceptions often lead to training interruptions, resulting in large loss of training progress, and the updated model parameters are stored after each iteration, which increases performance overhead and reduces training efficiency.

Method used

During the calculation of the adjustment parameter of distributed training, the model parameters of the deep learning model are stored according to the preset target storage parameters, and the target storage parameters are determined in combination with the model specification information of the deep learning model and the preset distributed training strategy, real-time storage of model parameters is realized, and training is quickly restored when training exceptions are trained.

Benefits of technology

It realizes rapid recovery of training under training exceptions, reduces training progress loss, and overlaps the stored procedures of model parameters with the adjustment parameter calculation process in distributed training, saving performance overhead and improving training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024118478_05062025_PF_FP_ABST
    Figure CN2024118478_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a deep learning model training method and a deep learning model training system. The deep learning model training method comprises: acquiring an initial deep learning model and a sample data set; and performing distributed training on the deep learning model on the basis of the sample data set according to a preset distributed training policy, and during the adjustment parameter calculation for distributed training, storing model parameters of the deep learning model on the basis of target storage parameters, wherein the target storage parameters are determined on the basis of model specification information of the deep learning model and the preset distributed training policy. The target storage parameters are determined on the basis of the model specification information of the deep learning model and the preset distributed training policy, which fully takes into account the iteration patterns of distributed training, and during the adjustment parameter calculation, the model parameters of the deep learning model are stored on the basis of the target storage parameters, so that high efficiency is achieved while enabling the training of the deep learning model to have high fault tolerance.
Need to check novelty before this filing date? Find Prior Art

Description

Deep learning model training method and deep learning model training system

[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on November 30, 2023, with application number 2023116363659 and application name “Deep Learning Model Training Method and Deep Learning Model Training System”, the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0002] The embodiments of the present disclosure relate to the field of deep learning technology, and in particular to a deep learning model training method and a deep learning model training system. Background Art

[0003] With the development of deep learning technology, large-scale deep learning models represented by large language models have been widely used in tasks in different scenarios.

[0004] Currently, distributed training has become the mainstream approach for efficiently training deep learning models. However, during distributed training, training anomalies are inevitable. These anomalies often lead to training interruptions, and the updated model parameters are not stored, resulting in significant training progress losses and insufficient stability in deep learning model training. Furthermore, storing the updated model parameters after each iteration during distributed training increases performance overhead and reduces training efficiency.

[0005] Summary of the Invention

[0006] In view of this, embodiments of the present disclosure provide a deep learning model training method. One or more embodiments of the present disclosure also relate to another deep learning model training method, a deep learning model training system, a deep learning model training apparatus, another deep learning model training apparatus, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.

[0007] According to a first aspect of an embodiment of the present disclosure, a deep learning model training method is provided, comprising:

[0008] Obtain an initial deep learning model and sample dataset;

[0009] According to the preset distributed training strategy, the deep learning model is distributedly trained based on the sample data set, and during the adjustment parameter calculation process of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.

[0010] According to a second aspect of an embodiment of the present disclosure, another deep learning model training method is provided, which is applied to a cloud-side device, where the cloud-side device includes multiple distributed nodes and a storage medium. The method includes:

[0011] Obtain an initial deep learning model and sample dataset;

[0012] According to a preset distributed training strategy, calling multiple distributed nodes, performing distributed training on the deep learning model based on the sample data set, and storing the model parameters of the deep learning model to a storage medium according to target storage parameters during the distributed training adjustment parameter calculation process, wherein the target storage parameters are determined based on model specification information of the deep learning model and the preset distributed training strategy;

[0013] When an abnormality in deep learning model training is identified, multiple distributed nodes are triggered to stop the distributed training of the deep learning model and determine the target model parameters currently stored in the storage medium;

[0014] Upon receiving a request to resume training, obtaining target model parameters from a storage medium;

[0015] Call multiple distributed nodes to perform distributed training of deep learning models based on target model parameter recovery.

[0016] According to a third aspect of an embodiment of the present disclosure, a deep learning model training system is provided, the system comprising a management and control unit and a plurality of distributed nodes, the plurality of distributed nodes including a first distributed node, the first distributed node being any one of the plurality of distributed nodes;

[0017] The control unit is used to obtain the initial deep learning model and sample data set, build multiple distributed data based on the deep learning model and sample data set according to the preset distributed training strategy, and distribute the multiple distributed data to each distributed node;

[0018] The first distributed node is used to perform distributed training on the deep learning model based on the sample data set; and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored therein according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.

[0019] According to a fourth aspect of an embodiment of the present disclosure, a deep learning model training apparatus is provided, comprising:

[0020] A first acquisition module is configured to acquire an initial deep learning model and a sample dataset;

[0021] The first training module is configured to perform distributed training on the deep learning model based on the sample data set according to a preset distributed training strategy, and store the model parameters of the deep learning model according to the target storage parameters during the adjustment parameter calculation process of the distributed training, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.

[0022] According to a fifth aspect of an embodiment of the present disclosure, another deep learning model training apparatus is provided, which is applied to a cloud-side device, wherein the cloud-side device includes multiple distributed nodes and a storage medium; the apparatus includes:

[0023] A second acquisition module is configured to acquire an initial deep learning model and a sample dataset;

[0024] a second training module configured to call multiple distributed nodes according to a preset distributed training strategy, perform distributed training on the deep learning model based on the sample data set, and store the model parameters of the deep learning model to a storage medium according to a target storage parameter during the distributed training adjustment parameter calculation process, wherein the target storage parameter is determined based on model specification information of the deep learning model and the preset distributed training strategy;

[0025] A stop module is configured to trigger the multiple distributed nodes to stop the distributed training of the deep learning model when an abnormality in the deep learning model training is identified, and to determine the target model parameters currently stored in the storage medium;

[0026] a parameter acquisition module, configured to acquire target model parameters from a storage medium upon receiving a request to resume training;

[0027] The recovery module is configured to call multiple distributed nodes to perform distributed training on the deep learning model based on target model parameter recovery.

[0028] According to a sixth aspect of an embodiment of the present disclosure, there is provided a computing device, including:

[0029] memory and processor;

[0030] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above method are implemented.

[0031] According to a seventh aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions, and the instructions implement the steps of the above method when executed by a processor.

[0032] According to an eighth aspect of an embodiment of the present disclosure, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.

[0033] In one embodiment of the present disclosure, an initial deep learning model and a sample data set are obtained; distributed training is performed on the deep learning model based on the sample data set according to a preset distributed training strategy, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy. The target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, fully considering the iterative law of distributed training. In the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, and the storage process of the model parameters and the adjustment parameter calculation process in the distributed training are overlapped, which fully saves performance overhead. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] FIG1 is a flowchart of a deep learning model training method provided by one embodiment of the present disclosure;

[0035] FIG2 is a system architecture diagram of a deep learning model training method provided by one embodiment of the present disclosure;

[0036] FIG3 is a flow chart of a deep learning model training method provided by one embodiment of the present disclosure;

[0037] FIG4 is a schematic diagram of a deep learning model training method provided by one embodiment of the present disclosure;

[0038] FIG5 is a flowchart of another deep learning model training method provided by one embodiment of the present disclosure;

[0039] FIG6 is a schematic diagram of the structure of a deep learning model training system provided by one embodiment of the present disclosure;

[0040] FIG7 is a schematic structural diagram of a deep learning model training device provided by one embodiment of the present disclosure;

[0041] FIG8 is a schematic structural diagram of another deep learning model training device provided by one embodiment of the present disclosure;

[0042] FIG9 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0043] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.

[0044] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.

[0045] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0046] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0047] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, which typically contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model (Foundation Model), which is pre-trained by using large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization ability, such as a large-scale language model (LLM) and a multi-modal pre-training model.

[0048] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0049] First, the terms involved in one or more embodiments of the present disclosure are explained.

[0050] Natural Language Processing (NLP) is a key area of ​​research in computer science and artificial intelligence. Its goal is to enable computers to understand and use human language to perform useful tasks. NLP is primarily used in areas such as machine translation, speech recognition, text analysis, and text question answering.

[0051] Hyperparameters: are fixed parameters calculated during model training. They can be understood as a strategy for updating model parameters and are used to control the update process of model parameters.

[0052] Gradient weight: This refers to the relationship between the gradient in a model and the model parameters. The gradient describes the direction and rate of change of the loss function with respect to the model parameters. The weight, on the other hand, refers to the degree of influence of the model parameters on the loss function. It determines how quickly the gradient descent algorithm updates the parameters; larger gradient weights result in faster parameter updates.

[0053] Grid search parameter tuning is a commonly used hyperparameter tuning strategy that involves searching through all possible hyperparameter combinations to find the target hyperparameter combination. Specifically, we first need to define a hyperparameter space that contains the possible value ranges for each hyperparameter. We then divide this hyperparameter space into a series of subspaces, each corresponding to a set of hyperparameter combinations. Next, we apply each hyperparameter combination in each subspace to the model and measure its performance. Ultimately, we find the target hyperparameter combination, which serves as the basis for updating the model parameters.

[0054] Bayesian optimization is a commonly used hyperparameter tuning strategy based on Bayes' theorem. Bayes' theorem is a theory in probability theory used to estimate the probability of an event and can be used to infer the target optimization value of hyperparameters. In Bayesian parameter tuning, we first define a prior distribution that represents our knowledge of the hyperparameters. We then combine the prior distribution with the observed experimental results to obtain the posterior distribution. Using this posterior distribution, we can estimate the target optimization value of the hyperparameters.

[0055] Mini-batch: In deep learning scenarios, by dividing the entire dataset into several smaller datasets and training on each smaller dataset at a time, this avoids the huge computational overhead of training all the data in the dataset at once. In addition, when performing gradient updates, the gradient direction does not differ significantly from the entire dataset, ensuring the training effect.

[0056] Graphics Processing Unit (GPU): A microprocessor used for graphics-related calculations. With the development of deep learning technology, its parallel structure can perform efficient matrix operations and is widely used as computing hardware for model training.

[0057] Tensor Processing Unit (TPU): A customized chip specifically for deep learning tasks that can provide higher efficiency and performance than GPU.

[0058] Field-Programmable Gate Array (FPGA): A programmable logic gate array that can be programmed to perform a variety of functions, including convolution operations in deep learning.

[0059] Application Specific Integrated Circuit (ASIC): An integrated circuit designed specifically for a specific application that can provide very high performance but at a higher cost.

[0060] A central processing unit (CPU) can also be used for model training, but it is generally not as efficient as a GPU or TPU.

[0061] Deep Self-Attention Model (Transformer Model): A deep learning architecture based on the attention mechanism for processing sequential data such as natural language.

[0062] Bidirectional Encoder Representations from Transformers (BERT): A special Transformer model trained using a bidirectional Transformer encoder and large-scale unlabeled text data. BERT's outstanding performance has made it a standard baseline for many natural language processing tasks.

[0063] Large Language Model (LLM): A deep learning model trained on a large corpus for natural language processing tasks. These models typically consist of a multi-layer neural network. Their input is a text sequence, which they generate. Their output is the generated text for a specific natural language processing task performed on that text sequence. Pre-training means that the model has been trained and learned to process large amounts of language data before a specific task. By pre-training the models, they can capture more complex linguistic and semantic rules, enabling them to excel in various natural language processing tasks and reducing the large-scale data requirements for specific tasks.

[0064] Distributed training: A method of training deep learning models using multiple distributed nodes, which can greatly increase training speed and reduce computation time.

[0065] Model Parallel (MP): A distributed training strategy that splits a model into multiple parts and deploys each part to different distributed nodes for training. This approach can effectively solve the problem of excessive parameters in the model.

[0066] Data Parallel (DP): A distributed training strategy that splits the entire sample set into multiple subsets and assigns each subset to a different distributed node for training. This technique can improve training efficiency and reduce memory usage.

[0067] Pipeline Parallel (PP): A distributed training strategy that lies between model parallelism and data parallelism. The core idea is to decompose large models into multiple layers and organize them into a pipeline to perform forward and backward propagation calculations, thereby reducing the graphics card memory usage and communication overhead.

[0068] Remote Direct Memory Access (RDMA): A technology used to directly read and write memory between remote computers, which can greatly improve the speed of data transmission in the network.

[0069] Network card: A hardware device that is mainly used to realize the physical connection between the computer and the network, and completes data transmission by sending and receiving data packets.

[0070] Peripheral Component Interconnect Express (PCIe): A physical connection structure between a group of nodes, usually consisting of a bus, that can be used to support data transmission between multiple devices.

[0071] Inter-GPU Express: A fast communication channel used to connect two or more GPUs, providing higher bandwidth and lower latency for data transmission between GPUs.

[0072] Communication channel: A path used to transmit data signals, which can be either a physical channel or a logical channel. A physical channel consists of the transmission medium and associated communication equipment, and is used to transmit the actual data signal. A logical channel, on the other hand, is a logical path established through intermediate nodes on top of the physical channel, essentially forming a logical path between the sender and receiver.

[0073] At present, the distributed training of deep learning models mainly builds multiple distributed data based on sample sets or model parameters of deep learning models, and then distributes multiple distributed data to different distributed nodes. Iterative training is performed on any distributed node. During any iterative training process, the language model training is completed according to the method of propagation calculation and parameter update.

[0074] However, once a training anomaly occurs, such as a hardware anomaly, a system anomaly, a network anomaly, or other unknown anomalies, distributed training needs to be re-executed without storing the updated model parameters. This is unbearable for deep learning models with large performance overhead. Although the updated model parameters can be stored at specific checkpoints (for example, after the parameters are updated in each iteration), when the deep learning model parameters are large, it is necessary to wait for the model parameters to be stored before continuing training. For deep learning models with model specifications reaching the migration level of tens of billions, such time overhead often reaches several minutes or even more than ten minutes, which determines that model parameters cannot be stored frequently. In this case, once a training anomaly occurs, the time overhead of resuming distributed training may reach several hours. Therefore, there is an urgent need for a deep learning model training method that has low performance overhead, pre-stores updated model parameters when a training anomaly occurs, and does not need to be recalculated after resuming distributed training, while having high stability and high efficiency.

[0075] In response to the above problems, the present disclosure provides a deep learning model training method. The present disclosure also involves another deep learning model training method, a deep learning model training system, a deep learning model training device, another deep learning model training device, a computing device, a computer-readable storage medium and a computer program, which are described in detail one by one in the following embodiments.

[0076] Referring to FIG1 , FIG1 shows a flowchart of a deep learning model training method provided by one embodiment of the present disclosure, including the following specific steps:

[0077] Step 102: Obtain an initial deep learning model and sample dataset.

[0078] The embodiments of the present disclosure are applied to a server with a distributed training function, on which multiple distributed nodes and storage media are deployed, and any distributed node includes computing hardware for model training, such as a GPU, NPU, FPGA, ASIC, or CPU.

[0079] A deep learning model refers to a type of machine learning model based on a deep neural network structure, which is widely used in multiple fields such as vision, speech, and natural language processing. Deep learning models have large-scale model parameters, and include but are not limited to: language processing models, image processing models, speech processing models, code processing models, etc. Taking the language processing model as an example, the language processing model can perform one or more natural language processing tasks, including but not limited to: machine translation tasks, speech recognition tasks, text analysis tasks, or text question and answer tasks. In terms of function, the language processing model can be regarded as: translation model, speech recognition model, text analysis model, text question and answer model, etc., without limitation here. In terms of model structure, the language processing model can be a Transformer model, a BERT model, a large language model, etc., without limitation here.

[0080] The sample data set is a collection of sample data for training a deep learning model, and the sample data set includes large-scale sample data. The sample data set can be a labeled sample data set or an unlabeled sample data set. Depending on the training requirements, the sample data can be data of different modalities. For example, if the deep learning model needs to be trained to become a model with text processing capabilities, the sample text of the text modality is used. For another example, if the deep learning model needs to be trained to become a model with audio processing capabilities, the sample audio of the audio modality is used. For another example, if the deep learning model needs to be trained to become a model with image processing capabilities, the sample image of the image modality is used. For another example, if the deep learning model needs to be trained to become a model with numerical processing capabilities, the sample numerical value of the numerical modality is used. The sample data set can be obtained from a sample database, such as an open source sample database, or it can be artificially constructed, such as generated using a generative model. It can also be obtained from a historical database, such as obtaining historical query text and historical answer text from a historical database to construct a sample text set. This is not limited here.

[0081] For example, a website for a large language model needs to add a virtual character dialogue feature. The initial large language model is trained to enable the resulting target large language model to have text question-answering capabilities, enabling it to perform text question-answering tasks. The initial large language model is obtained from the model library, and a sample text set consisting of 10,000,000 sample question-answer text pairs is obtained from an open-source sample database.

[0082] Obtain the initial deep learning model and sample dataset. This lays the foundation for the model and sample data for subsequent distributed training.

[0083] Step 104: According to the preset distributed training strategy, the deep learning model is distributedly trained based on the sample data set, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.

[0084] Distributed training is a method for training deep learning models using multiple distributed nodes. Model training is iterative, and each iteration includes parameter adjustment calculations and model parameter updates. Adjustment parameters are hyperparameters used to adjust model parameters, and parameter adjustment calculations are the calculations of these hyperparameters, including but not limited to propagation calculations, probability calculations (updating model parameters through Bayesian optimization), grid calculations (updating model parameters through grid search), and diffusion inference processes (using a diffusion model for forward and backward diffusion).

[0085] Distributed training strategies employ a distributed computing framework to split large-scale deep learning model training tasks into multiple smaller tasks, distributing them across multiple distributed nodes for execution. These strategies include, but are not limited to, data parallelism, model parallelism, and pipeline parallelism. Distributed training strategies are pre-defined based on the model properties of the deep learning model, the sample dataset, and / or the training task. For example, if the layers of a deep learning model have sequential execution constraints, making it difficult to split the training, a data parallel strategy may be employed. Another example is if the sample dataset is too large, a data parallel strategy may be employed. Furthermore, in large-scale text classification training tasks, adopting a data parallel strategy may cause network bottlenecks, while adopting a model parallel strategy may result in excessive complexity, thus adopting a pipeline parallel strategy. Different distributed training strategies correspond to different iterative characteristics of model training. For example, with a data parallel strategy, the amount of sample data on each distributed node is small, while the model parameters are large, resulting in high time overhead for parameter adjustment calculations and a low number of iterations. With a model parallel strategy, the amount of sample data on each distributed node is large, while the model parameters are small, resulting in low time overhead for parameter adjustment calculations and a high number of iterations. Different iteration characteristics determine different time costs for adjusting parameter calculations and storage times. Therefore, it is necessary to fine-tune the target storage parameters and overlap the model parameter storage process with the parameter adjustment calculation process in distributed training.

[0086] Model specification information refers to the model parameter specifications of a deep learning model, including but not limited to model parameter counts and benchmark models. Model parameter counts refer to the parameter specifications of model parameters. For example, the model parameter count of a deep learning model may be in the tens of billions. A benchmark model refers to a specific deep learning model architecture. For example, in natural language processing, benchmark models for deep learning models include the Transformer model, the BERT model, and the Large Language Model.

[0087] In the disclosed embodiment, the storage is performed synchronously with the parameter adjustment calculation process, that is, the storage process and the parameter adjustment calculation process have a time overhead that is not much different. For example, the time overhead of the parameter adjustment calculation process is T, and the time overhead of the storage process is T'. If T is less than T', even during the model parameter update process, the storage is still being performed. Once a training anomaly occurs, it is difficult to resume model training. For details, see the following description. Taking propagation calculation as an example, the storage process can be performed during the forward propagation calculation process, the reverse propagation calculation process, or both the forward propagation calculation process and the reverse propagation calculation process, without limitation here.

[0088] Target storage parameters are configuration parameters for storing model parameters. These parameters are determined based on the deep learning model's specifications and the pre-set distributed training strategy, and include, but are not limited to, target storage model parameter specifications, target storage frequency, target storage read / write speed, and target storage bandwidth. Once a distributed training strategy is determined and executed, the time required to adjust parameter calculations is difficult to change. To ensure that the model parameter storage process overlaps with the parameter adjustment calculation process during distributed training, target storage parameters are used to control the time required to store model parameters.

[0089] It should be noted that if the model parameters are stored during the model parameter update process, the updated model parameters and the unupdated model parameters will be mixed and stored. Once a training anomaly occurs, the model training cannot be resumed by storing the mixed updated model parameters and the unupdated model parameters. For example, after completing the i-1th iteration training, the current model parameters are θ i-1 , perform the i-th iteration training, and in the process of adjusting the parameter calculation, the current model parameter θ i-1 After the storage is completed, if a training anomaly occurs, the model parameters θ can be obtained i-1 Re-execute the i-th iteration training. If during the model parameter update process, the current model parameter θ i-1 For storage, some updated model parameters θ will be introduced i , and thus it is impossible to re-execute the i-th iteration training when a training anomaly occurs.

[0090] According to the preset distributed training strategy, the deep learning model is distributedly trained based on the sample data set. The specific method is: according to the preset distributed training strategy, multiple distributed data are constructed, the multiple distributed data are distributed to distributed nodes, and iterative training is performed. Among them, the iterative training includes adjusting parameter calculation and model parameter update.

[0091] The model parameters of the deep learning model are stored according to the target storage parameters. Specifically, the model parameters of the deep learning model are stored in a storage medium according to the target storage parameters. The storage medium is the hardware device that stores the model parameters, including but not limited to: computing hardware cache (GPU cache, NPU cache, FPGA cache, ASIC cache and CPU cache), memory, hard disk and distributed persistent storage array.

[0092] The target storage parameters are determined based on the model specifications of the deep learning model and the preset distributed training strategy. Specifically, a time cost analysis is performed based on the model specifications of the deep learning model and the preset distributed training strategy to obtain the target storage parameters. The time cost analysis involves analyzing the time cost of adjusting parameter calculations and storing model parameters during iterative training, thereby determining the target storage parameters that can achieve storage.

[0093] For example, the preset distributed training strategy is a data parallel strategy. The number of model parameters based on the large language model is 10 13 Using a level and data parallel strategy, a time cost analysis is performed to determine the time cost t1 of the forward propagation computation process and the time cost t2 of the backward propagation computation process for a single iteration. Based on these time costs and the number of model parameters, the target storage model parameter specifications P1 and P2 are obtained. Using the data parallel strategy, 10,000,000 sample question-answer text pairs in the sample text set are partitioned. 64 distributed data sets are constructed and distributed to 64 distributed nodes, each equipped with a GPU. On each distributed node, each distributed data set is divided into 16 small batches, and iterative training of 16 small batches is performed using the GPU. During the forward and backward propagation computations of each iterative training, the model parameters θ of the large language model are sequentially stored in the GPU cache, memory, and hard disk of the distributed nodes according to the target storage model parameter specifications P1 and P2. Following this strategy, after completing distributed training, a target large language model with text question-answering functionality is obtained. This target large language model is deployed on the cloud-side device of the large language model's website, providing users with virtual character dialogue functionality.

[0094] In the disclosed embodiment, an initial deep learning model and a sample data set are obtained; distributed training is performed on the deep learning model based on the sample data set according to a preset distributed training strategy, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy. The target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, fully considering the iterative law of distributed training. In the process of calculating the adjustment parameters of distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, and the storage process of the model parameters and the adjustment parameter calculation process in distributed training are overlapped, which fully saves performance overhead. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency.

[0095] In an optional embodiment of the present disclosure, step 104 performs distributed training on the deep learning model based on the sample dataset according to a preset distributed training strategy, including the following specific steps:

[0096] Construct multiple distributed data based on the deep learning model and the sample data set according to the preset distributed training strategy, wherein the preset distributed training strategy includes a model parallel training strategy or a data parallel training strategy;

[0097] Distribute multiple distributed data to each distributed node;

[0098] Performing propagation calculation based on the distributed data on a first distributed node to obtain a gradient weight, wherein the first distributed node is any one of the multiple distributed nodes;

[0099] Based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated, and when the preset training end conditions are met, a trained deep learning model is obtained.

[0100] Distributed data refers to the portion of data distributed and executed on distributed nodes. If the entire model training is considered a large-scale training task, distributed data is the task data of the smaller training tasks resulting from splitting that large-scale training task. Different distributed training strategies result in different distributed data. For example, using a model parallel strategy can split the model parameters of a deep learning model into multiple components to construct distributed data. Another example is using a data parallel strategy to split a sample dataset into multiple components to construct distributed data.

[0101] Propagation calculation is the process of determining the gradient weight based on sample data using a deep learning model, including the forward propagation calculation process and the back propagation calculation process. Forward propagation calculation is the process of inputting sample data into the deep learning model and outputting predicted data. Back propagation calculation is the process of determining the loss value through predicted data, and then inputting the deep learning model in reverse to determine the gradient weights of each model layer. Model parameter update is the process of adjusting the model parameters of each model layer based on the gradient weights. For example, in the forward propagation calculation process, sample data X is input into the large language model and predicted data Z is output. In the back propagation calculation process, the loss value is determined based on the predicted data Z, and the large language model is input in reverse to determine the gradient weights of the n model layers of the large language model. for: During the model parameter update process, based on the gradient weight Adjust the model parameters of n model layers in a large language model.

[0102] Gradient weights are used to assign different gradient weights to the model parameter updates of each model layer during the backward propagation process, thereby controlling the update amplitude and linear speed of the model parameter updates. The larger the parameter update value for each model layer, the faster the parameter update, and vice versa.

[0103] The preset training end conditions are pre-set judgment conditions for stopping training, including but not limited to: a preset number of iterations, a preset loss value threshold, a preset training time and a preset model convergence condition.

[0104] According to the preset distributed training strategy, multiple distributed data are constructed based on the deep learning model and the sample data set. The specific method is: according to the preset distributed training strategy, the model parameters and / or the sample data set of the deep learning model are divided to obtain multiple distributed data.

[0105] Propagation calculation is performed based on distributed data to obtain gradient weights. The specific method is: sample data in the distributed data is input into the deep learning model, forward propagation calculation is performed to obtain predicted data, loss value is determined based on sample data and predicted data, loss value is reversely input into the deep learning model, back propagation calculation is performed to obtain gradient weights.

[0106] Based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated. Specifically, based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated through the gradient update method.

[0107] For example, according to the data parallel strategy, 10,000,000 sample question-answer text pairs in the sample text set are divided to obtain 64 distributed data. The 64 distributed data are distributed to 64 distributed nodes, each of which is deployed with a GPU. On any distributed node, any distributed data is divided into 16 small batches. During any iterative training process, the sample question text X in the small batch of distributed data is input into the large language model, and the forward propagation calculation is performed to obtain the predicted answer text Z′. Based on the sample answer text Z and the predicted answer text, the loss value Loss is determined, and the loss value is reversely input into the large language model, and the backpropagation calculation is performed to obtain the gradient weight. Based on the gradient weights on each distributed node, the model parameters θ of the large language model are updated through the gradient update method. After completing the distributed data training of all small batches, the trained target large language model is obtained. The target large language model has text question and answer capabilities.

[0108] According to the preset distributed training strategy, multiple distributed data are constructed based on the deep learning model and the sample data set, wherein the preset distributed training strategy includes a model parallel training strategy or a data parallel training strategy; the multiple distributed data are distributed to each distributed node; on the first distributed node, propagation calculation is performed based on the distributed data to obtain gradient weights, wherein the first distributed node is any one of the multiple distributed nodes; based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated, and when the preset training end conditions are met, a trained deep learning model is obtained. According to the preset distributed training strategy, multiple distributed data are constructed, distributed training is performed on multiple distributed nodes, gradient weights are obtained, and the model parameters of the deep learning model are updated, thereby improving the efficiency of model training.

[0109] In an optional embodiment of the present disclosure, the distributed data includes multiple batches of distributed data;

[0110] On the first distributed node, a propagation calculation is performed based on the distributed data to obtain the gradient weight, including the following specific steps:

[0111] On the first distributed node, propagation calculation is performed based on the distributed data of the current batch to obtain the gradient weight;

[0112] Correspondingly, after updating the model parameters of the deep learning model based on the gradient weights on each distributed node, the following specific steps are also included:

[0113] Update the distributed data of the current batch, return to the step of performing propagation calculation on the first distributed node based on the distributed data of the current batch, and obtain the gradient weight.

[0114] The distributed data of the current batch is the distributed data of the batch used to train the deep learning model in the current iterative training process. For example, on any distributed node, the distributed data is divided into 16 batches, and 16 iterative trainings are required. One iterative training process includes the forward propagation calculation process, the backpropagation calculation process, and the parameter update process. Optionally, before the parameter update process, a communication process (the process of integrating gradient weights) is also included. The current iterative training process is the i-th iterative training process, the distributed data of the current batch is the distributed data of the i-th batch, and the current model parameters are the model parameters θ after the i-1-th update. i-1 Correspondingly, the distributed data of the current batch is updated, that is, the distributed data of the i-th batch is updated to the distributed data of the i+1-th batch.

[0115] Propagation calculation is performed based on the distributed data of the current batch to obtain the gradient weight. The specific method is as follows: the sample data in the distributed data of the current batch is input into the deep learning model, and the forward propagation calculation is performed to obtain the predicted data. Based on the sample data and the predicted data, the loss value is determined, and the loss value is reversely input into the deep learning model, and the back propagation calculation is performed to obtain the gradient weight.

[0116] For example, the sample question text X in the distributed data of batch i is i Input large language model (the current model parameter is the model parameter θ after the i-1th update) i-1 ), perform forward propagation calculation, obtain the predicted data prediction answer text Z′ i , based on the sample answer text Z i And predict the answer text, determine the loss value Loss, input the loss value back into the large language model, perform back propagation calculation, and obtain the gradient weight Based on the gradient weights on each distributed node, the model parameters of the large language model are updated from θ i-1 Update to θ i , update the current batch of distributed data from batch i to batch i+1, and continue to execute the sample question text X in the distributed data of batch i+1 i+1 The steps of inputting a large language model are as follows: After completing the distributed data training of all small batches, a trained target large language model is obtained, and the target large language model has a text question-answering function.

[0117] On the first distributed node, propagation calculations are performed based on the current batch of distributed data to obtain gradient weights. The current batch of distributed data is updated, and the process returns to the first distributed node, where propagation calculations are performed based on the current batch of distributed data to obtain gradient weights. By dividing the distributed data into multiple batches and iteratively training the deep learning model using these batches of distributed data, this avoids the huge computational overhead of requiring all distributed data to participate in distributed training at once, which can lead to training bottlenecks and ensures effective training.

[0118] In an optional embodiment of the present disclosure, before performing propagation calculation on the distributed data of the current batch and obtaining the gradient weights on the first distributed node, the following specific steps are further included:

[0119] For each batch of distributed data, based on the model specifications of the deep learning model and the preset distributed training strategy, the number of propagation calculations and time overhead are predicted;

[0120] Based on the number of propagation calculations and the time overhead, the target storage parameters corresponding to the propagation calculation process are determined.

[0121] In the embodiment of the present disclosure, the number of forward and backward calculations and the corresponding time overhead required for each batch of distributed data in the training process of the deep learning model under different distributed training strategies are also different. For example, when a data parallel strategy is adopted, the amount of sample data on each distributed node is small, while the model parameter specifications are large, the time overhead of a single propagation calculation is high, and the number of times is small. When a model parallel strategy is adopted, the amount of sample data on each distributed node is large, while the model parameter specifications are small, the time overhead of a single propagation calculation is low, and the number of times is large. It is necessary to finely determine the target storage parameters based on the number and time overhead of propagation calculations, and overlap the storage process of the model parameters in each iterative training process with the propagation calculation process in distributed training.

[0122] For each batch of distributed data, based on the model specifications of the deep learning model and the preset distributed training strategy, the number of propagation calculations and the time overhead are predicted. Specifically, the predictions are obtained through prediction algorithms, such as the torch.distributed module, the torch.profiler tool, and the tf.data API interface.

[0123] Based on the number of propagation calculations and the time overhead, the target storage parameters corresponding to the propagation calculation process are determined. The specific method is: with the time overhead of the model parameter storage process not exceeding the time overhead of the propagation calculation process as the goal, based on the number of propagation calculations and the time overhead, the target storage parameters corresponding to the propagation calculation process are determined.

[0124] For example, through the torch.distributed module, for each batch of distributed data, the number of model parameters based on the large language model is 10 13 Level and data parallel strategy, predicting the number of propagation calculations to be 16 times, the time overhead of the forward propagation calculation process is t1, and the time overhead of the backward propagation calculation process is t2. With the goal that the time overhead T' of the model parameter storage process does not exceed the time overhead T=16*(t1+t2) of the propagation calculation process, the target storage model parameter specifications corresponding to the propagation calculation process are determined to be P1 and P2.

[0125] For each batch of distributed data, the number and time of propagation calculations are predicted based on the deep learning model's specifications and the pre-set distributed training strategy. Based on this number and time, the target storage parameters for the propagation calculation process are determined. This ensures the feasibility of subsequent storage of model parameters and more accurately overlaps the storage process of model parameters with the propagation calculation process during distributed training.

[0126] In an optional embodiment of the present disclosure, before updating the model parameters of the deep learning model based on the gradient weights on each distributed node, the following specific steps are also included:

[0127] The gradient weights on each distributed node are integrated through the communication channels between the distributed nodes.

[0128] The communication channels between distributed nodes are the data transmission channels between them. They can be physical connections, such as network cards, PCIe topologies, or optical fibers, or virtual connections, such as high-speed channels between GPUs or between distributed nodes. Communication channels can be implemented using RDMA technology, which increases transmission speed.

[0129] For example, a high-speed channel is built through the network cards on 64 distributed nodes, and the gradient weights on each distributed node are integrated using RDMA technology. Based on the gradient weights on each distributed node, the model parameters θ of the large language model are updated through the gradient update method.

[0130] The communication channels between distributed nodes are used to integrate the gradient weights of each distributed node. The communication process of iterative training is centralized through the communication channels, which improves the efficiency and stability of model training.

[0131] In an optional embodiment of the present disclosure, any distributed node includes a first communication channel connected to a storage medium and a second communication channel connected to other distributed nodes;

[0132] In step 104, the model parameters of the deep learning model are stored according to the target storage parameters, including the following specific steps:

[0133] Storing the model parameters of the deep learning model to a storage medium via the first communication channel according to the target storage parameters;

[0134] Correspondingly, the gradient weights on each distributed node are integrated through the communication channel between the distributed nodes, including the following specific steps:

[0135] The gradient weights on each distributed node are integrated through the second communication channel.

[0136] In the embodiment of the present disclosure, the storage medium is independent of the distributed nodes, avoiding anomalies on any distributed node that may cause data loss and ensuring the reliability of the entire distributed system. Therefore, the distributed nodes need to establish a communication channel with the storage medium for data storage. However, the resource performance on the distributed nodes is limited, and the communication channel is needed in the process of storing model parameters. In the communication process, that is, integrating the gradient weights on each distributed node, the communication channel is also needed. This is difficult to achieve for large-scale deep learning models. Therefore, it is necessary to distinguish the communication channels of the two processes and isolate the two data transmission processes to avoid introducing additional collective communications and ensure that no interference is caused to the distributed training.

[0137] The first communication channel connecting the storage medium is a communication channel used for the model parameter storage process, including physical connections and virtual connections between distributed nodes and storage media, such as network cards, PCIe topology, optical fibers, read and write channels (data buses) between distributed nodes and storage media, etc.

[0138] The second communication channel connecting to other distributed nodes is used for communication. It includes physical and virtual connections between distributed nodes and storage media, such as network cards, PCIe topologies, optical fibers, high-speed channels between distributed nodes and GPUs, and high-speed channels between distributed nodes. This second communication channel can be implemented using RDMA technology, which improves transmission speed.

[0139] Exemplarily, the first communication channel is the physical connection between the network cards on the 64 distributed nodes and the distributed persistent storage array. According to the target storage model parameter specifications P1 and P2, the model parameters θ of the large language model are stored in the distributed persistent storage array through the first communication channel. The second communication channel is the physical connection between the network cards on the 64 distributed nodes, and the PCIe topology with GPUs plugged into each distributed node, with virtual connections between high-speed channels between GPUs and high-speed channels between distributed nodes. Through the second communication channel, using RDMA technology, the gradient weights on each distributed node are integrated Based on the gradient weights on each distributed node, the model parameters θ of the large language model are updated through the gradient update method.

[0140] Based on the target storage parameters, the deep learning model parameters are stored to the storage medium via the first communication channel. The gradient weights on each distributed node are integrated via the second communication channel. This separates the two data transmission processes, model parameter storage and communication, to avoid introducing additional collective communications and ensure no interference with distributed training. This improves the reliability of distributed training and enhances the effectiveness of model training.

[0141] In an optional embodiment of the present disclosure, the propagation calculation includes forward propagation calculation and backward propagation calculation;

[0142] In the propagation calculation process of the distributed training in step 104, before storing the model parameters of the deep learning model according to the target storage parameters, the following specific steps are also included:

[0143] Based on the model specification information of the deep learning model and the preset distributed training strategy, determine the first target storage parameters corresponding to the forward propagation calculation process and the second target storage parameters corresponding to the backward propagation calculation process;

[0144] Correspondingly, in step 104, during the propagation calculation process of distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, including the following specific steps:

[0145] During the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameters;

[0146] During the back propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameters.

[0147] The propagation calculation process includes the forward propagation calculation process and the backward propagation calculation process. The time overhead required for the two is different. Therefore, it is necessary to make a detailed distinction and determine the target storage parameters in the forward propagation calculation process and the target storage parameters in the backward propagation calculation process.

[0148] The first target storage parameters are configuration parameters for storing model parameters during the forward propagation calculation process. The first target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, including but not limited to: first target storage model parameter specifications, first target storage frequency, first target storage read and write speed, first target storage bandwidth, etc.

[0149] The second target storage parameters are configuration parameters for storing model parameters during the back-propagation calculation process. The second target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, including but not limited to: second target storage model parameter specifications, second target storage frequency, second target storage read and write speed, second target storage bandwidth, etc.

[0150] Based on the model specification information of the deep learning model and the preset distributed training strategy, determine the first target storage parameter corresponding to the forward propagation calculation process and the second target storage parameter corresponding to the backward propagation calculation process. The specific method is: for each batch of distributed data, based on the model specification information of the deep learning model and the preset distributed training strategy, predict the number and time overhead of the forward propagation calculation and the backward propagation calculation, and based on the number and time overhead of the forward propagation calculation and the backward propagation calculation, determine the first target storage parameter corresponding to the forward propagation calculation and the second target storage parameter corresponding to the backward propagation calculation process.

[0151] During the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameters. Specifically, during the forward propagation calculation process, the model parameters of the deep learning model are stored in the storage medium according to the first target storage parameters.

[0152] During the back propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameters. Specifically, during the back propagation calculation process, the model parameters of the deep learning model are stored in the storage medium according to the second target storage parameters.

[0153] For example, through the multi-process parallel communication module, for each batch of distributed data, the model parameter amount based on the large language model is 10 13Based on the level and data parallelism strategy, the number of forward propagation calculations is predicted to be 16 with a time overhead of t1, and the number of backward propagation calculations is predicted to be 16 with a time overhead of t2. With the goal that the time overhead of the model parameter storage process T' does not exceed the time overhead of the propagation calculation process T = 16*(t1+t2), the first target storage model parameter specification corresponding to the forward propagation calculation process is determined to be P1, and the second target storage model parameter specification corresponding to the backward propagation calculation process is determined to be P2. During the forward propagation calculation process, the model parameters θ of the large language model are stored in the GPU cache, memory, and distributed persistent storage array of the distributed nodes according to the first target storage model parameter specification. During the backward propagation calculation process, the model parameters θ of the large language model are stored in the GPU cache, memory, and distributed persistent storage array of the distributed nodes according to the second target storage model parameter specification.

[0154] Based on the model specification information of the deep learning model and the preset distributed training strategy, the first target storage parameter corresponding to the forward propagation calculation process and the second target storage parameter corresponding to the backward propagation calculation process are determined. During the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameter; during the backward propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameter. The target storage parameters corresponding to forward propagation and backward propagation are divided more finely, and the storage process of model parameters is overlapped with the forward propagation calculation process and backward propagation calculation process in distributed training. This completes the real-time storage of deep learning model parameters with near-zero performance overhead, making deep learning model training highly fault-tolerant and efficient.

[0155] In an optional embodiment of the present disclosure, storing the model parameters of the deep learning model according to the target storage parameters in step 104 includes the following specific steps:

[0156] Model parameters of the deep learning model are stored in multiple storage media according to target storage parameters and storage performance priorities of the multiple storage media.

[0157] Different storage media have different storage performance, including but not limited to read / write speed and persistence. For example, the read / write speed of memory is higher than that of a hard disk, but the persistence of memory is lower than that of a hard disk. Compared to a hard disk, memory has a lower tolerance for abnormalities.

[0158] According to the target storage parameters and the storage performance priorities of the multiple storage media, the model parameters of the deep learning model are stored in multiple storage media. The specific method is as follows: according to the target storage parameters and the storage performance priorities of the multiple storage media, a storage strategy is determined, and according to the storage strategy, the model parameters of the deep learning model are stored in multiple storage media. The embodiment of the present disclosure provides a storage strategy: establishing a storage medium hierarchy: from top to bottom, they are GPU cache, memory and hard disk, 1. Use higher-level storage media with faster read and write speeds as much as possible to maximize the parameter specification and abnormal recovery capabilities of the model parameters; 2. Even if the upper-level storage medium is unavailable, the lower-level, more persistent storage medium can still be relied on to ensure the persistent storage of the current model parameters; 3. Overhead is saved through asynchronous execution between the upper and lower layers.

[0159] Optionally, according to the target storage parameters, the model parameters of the deep learning model are stored in a multi-level storage medium, including the following specific steps: according to the target storage parameters, the model parameters of the deep learning model are stored in a first storage medium and a second storage medium respectively, wherein the read and write speed of the first storage medium is faster than that of the second storage medium.

[0160] Exemplarily, according to the target storage model parameter specifications P1 and P2, the model parameters θ of the large language model are stored in the GPU cache, memory and hard disk of the distributed nodes respectively.

[0161] Optionally, according to the target storage parameters, the model parameters of the deep learning model are stored in a multi-level storage medium, including the following specific steps: according to the target storage parameters, the model parameters of the deep learning model are stored in a first storage medium, so that the model parameters are transferred from the first storage medium to the second storage medium, wherein the reading and writing speed of the first storage medium is faster than that of the second storage medium.

[0162] Exemplarily, according to the target storage model parameter specifications P1 and P2, the model parameters θ of the large language model are stored in the GPU cache, so that the model parameters are transferred from the GPU cache to the memory, so that the model parameters are transferred from the memory to the hard disk.

[0163] The model parameters of the deep learning model are stored across multiple storage media according to the target storage parameters and the storage performance priorities of the multiple storage media. This fully utilizes the storage performance of different storage media, enabling high-frequency model parameter storage and storing the current model parameters, saving overhead.

[0164] In an optional embodiment of the present disclosure, the method further includes the following specific steps:

[0165] Upon receiving a resume training request, obtaining stored target model parameters, wherein the resume training request is generated after determining that the distributed training has recovered from a training anomaly, and the target model parameters are model parameters stored before the training anomaly occurs;

[0166] Resume distributed training of deep learning models based on target model parameters.

[0167] Training exceptions are abnormalities that occur during model training. These abnormalities may prevent model training from running properly or affect the performance and accuracy of the trained model. These include, but are not limited to, hardware, system, network, or other unknown anomalies. Training exception recovery involves taking appropriate measures to resume model training after a training exception occurs. A training resume request is a request to resume model training.

[0168] The target model parameters are the model parameters stored before the training anomaly occurs. For example, after completing the i-1th iteration training, the target model parameters are θ i-1 , perform the i-th iteration training, during the propagation calculation process, the target model parameter θ i-1 After the storage is completed, if a training anomaly occurs, the model parameters θ can be obtained i-1 Re-execute the i-th iteration training.

[0169] The stored target model parameters are obtained by: obtaining the stored target model parameters from a storage medium.

[0170] Exemplarily, upon receiving a request to resume training, the stored target model parameters θ are retrieved from the hard disk. Based on the target model parameters θ, distributed training of the large language model is resumed to obtain a trained target large language model having a text question-answering function.

[0171] Upon receiving a training resumption request, the stored target model parameters are retrieved. The training resumption request is generated after determining that a distributed training anomaly has been recovered, and the target model parameters are the model parameters stored before the training anomaly occurred. Distributed training of the deep learning model is resumed based on the target model parameters. This increases the fault tolerance of deep learning model training, avoids retraining the deep learning model, maintains stability, ensures training efficiency, and reduces training costs.

[0172] FIG2 shows a system architecture diagram of a deep learning model training method provided by one embodiment of the present disclosure, as shown in FIG2 :

[0173] The system architecture includes multiple layers of storage media, from top to bottom: the GPU cache, memory, and hard disk. The multi-layer storage media has increasingly higher persistence from top to bottom, and increasingly faster read and write speeds from bottom to top. During distributed training, the model parameters of the deep learning model are stored from the GPU to the GPU cache for non-persistent high-speed reading and writing. The model parameters of the deep learning model are stored from the GPU to the memory for non-persistent high-speed reading and writing. The model parameters of the deep learning model are transferred from the memory to the hard disk for persistent storage. When training needs to be resumed, the model parameters of the deep learning model are retrieved from the multi-layer storage media and the model training is performed again.

[0174] FIG3 shows a flow chart of a deep learning model training method provided by one embodiment of the present disclosure, as shown in FIG3 :

[0175] Currently, each iteration of model training includes propagation calculation, parameter update, and parameter storage. During iteration (i-1), propagation calculation and parameter update are completed, and the updated model parameters are stored. The i-th iteration begins, and the propagation calculation and parameter update are also completed, and the updated model parameters are stored.

[0176] In the embodiment of FIG1 of the present disclosure, each iteration includes propagation calculation, communication integration, and parameter update processes. During the propagation calculation process of the i-1th iteration, the model parameters are stored, followed by the communication integration and parameter update processes. The i-th iteration begins. During the propagation calculation process of the i-th iteration, the model parameters are stored, followed by the communication integration and parameter update processes.

[0177] Compared with the two, it saves time and improves the efficiency of model training.

[0178] FIG4 shows a schematic diagram of a deep learning model training method provided by an embodiment of the present disclosure, as shown in FIG4 :

[0179] The system includes multiple distributed nodes (two distributed nodes in the figure) and storage media (a distributed persistent storage array in the figure). Each distributed node includes memory, multiple graphics processing units, a first network interface card (NIC), and a second network interface card (NIC). High-speed communication channels are established between the NICs in the distributed nodes. A first communication channel is established between the first NIC on the distributed node and the distributed storage node. A second communication channel is established between the second NICs in the distributed nodes.

[0180] Referring to FIG5 , FIG5 shows a flowchart of another deep learning model training method provided by one embodiment of the present disclosure, which is applied to a cloud-side device, the cloud-side device including multiple distributed nodes and a storage medium; the method includes the following specific steps:

[0181] Step 502: Obtain an initial deep learning model and sample dataset.

[0182] Step 504: According to the preset distributed training strategy, multiple distributed nodes are called to perform distributed training on the deep learning model based on the sample data set, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored in the storage medium according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.

[0183] Step 506: When an abnormality in deep learning model training is identified, trigger multiple distributed nodes to stop distributed training of the deep learning model and determine the target model parameters currently stored in the storage medium.

[0184] Step 508: When a request to resume training is received, the target model parameters are obtained from the storage medium.

[0185] Step 510: Call multiple distributed nodes to perform distributed training on the deep learning model based on target model parameter recovery.

[0186] The embodiments of the present disclosure are applied to a cloud-side device with a distributed training function. The cloud-side device is a network cloud device, a virtual device, and is composed of multiple distributed nodes and storage media. Any distributed node includes computing hardware for model training, such as GPU, NPU, FPGA, ASIC or CPU.

[0187] The embodiment of the present disclosure and the embodiment of the specification of FIG. 1 are based on the same inventive concept. The specific methods of steps 502 , 504 , 508 and 510 have been described in detail in the embodiment of the specification of FIG. 1 and will not be repeated here.

[0188] For example, on a website of a large language model, it is currently necessary to add a virtual character dialogue function. By training the initial large language model, the target large language model obtained by training has a text question-answering function and can perform text question-answering tasks. The initial large language model is obtained from the model library, and the sample text set is obtained from the open source sample database. The sample text set includes 10,000,000 sample question-answer text pairs. The preset distributed training strategy is the data parallel strategy. The model parameter quantity based on the large language model is 10 13Using a level and data parallel strategy, we perform a time cost analysis, determining the time cost of the forward propagation process for a single iteration as t1 and the time cost of the backward propagation process as t2. Based on this time cost and the number of model parameters, we obtain the target storage model parameter specifications as P1 and P2. Using the data parallel strategy, we partition the 10,000,000 sample question-answer pairs in the sample text set. We construct 64 distributed data sets and distribute them across 64 distributed nodes, each equipped with a GPU. On each distributed node, we partition each distributed data set into 16 small batches, and use the GPU to perform iterative training on each of these 16 small batches. During the forward propagation and backpropagation calculations of each iterative training, the model parameters θ of the large language model are sequentially stored in the GPU cache, memory, and hard disk of the distributed nodes, according to the target storage model parameter specifications P1 and P2. If an anomaly in large language model training is identified, the 64 distributed nodes are triggered to stop distributed training of the large language model and determine the target model parameters θ currently stored on the hard disk. Upon receiving a request to resume training, the stored target model parameters θ are retrieved from the hard disk. Based on the target model parameters θ, distributed training of the large language model is resumed to obtain a trained target large language model. The target large language model has text question-answering capabilities and is deployed on the cloud-side device of the large language model's website to provide users with virtual character dialogue capabilities.

[0189] In an embodiment of the present disclosure, an initial deep learning model and a sample data set are obtained; according to a preset distributed training strategy, multiple distributed nodes are called to perform distributed training on the deep learning model based on the sample data set, and during the calculation of adjustment parameters for the distributed training, the model parameters of the deep learning model are stored in a storage medium according to target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy; when an abnormality in the deep learning model training is identified, multiple distributed nodes are triggered to stop the distributed training of the deep learning model, and the target model parameters currently stored in the storage medium are determined; when a request to resume training is received, the target model parameters are obtained from the storage medium; and multiple distributed nodes are called to resume distributed training of the deep learning model based on the target model parameters. The target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, and the iterative law of distributed training is fully considered. In the process of calculating the adjustment parameters of distributed training, the model parameters of the deep learning model are stored in the storage medium according to the target storage parameters, so that the storage process of the model parameters and the adjustment parameter calculation process in distributed training are overlapped, which fully saves performance overhead. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency. When the deep learning model training anomaly is identified, the stored target model parameters are obtained from the storage medium and the distributed training of the deep learning model is restored, which increases the fault tolerance of the deep learning model training and avoids re-training of the deep learning model. It ensures stability while ensuring training efficiency and reduces training costs.

[0190] In an optional embodiment of the present disclosure, the storage medium includes a plurality of storage media with different storage performances;

[0191] In step 504, during the distributed training adjustment parameter calculation process, the model parameters of the deep learning model are stored in the storage medium according to the target storage parameters, including the following specific steps:

[0192] Storing model parameters of the deep learning model in multiple storage media according to target storage parameters and storage performance priorities of the multiple storage media;

[0193] Correspondingly, obtaining the target model parameters from the storage medium in step 508 includes the following specific steps:

[0194] Acquiring target model parameters from a first storage medium;

[0195] If not obtained, the target model parameters are obtained from the second storage medium, wherein the storage performance priority of the first storage medium is higher than that of the second storage medium.

[0196] The steps of storing the model parameters of the deep learning model in multiple storage media according to the target storage parameters and the storage performance priorities of multiple storage media have been described in detail in the embodiment of the specification of Figure 1 above and will not be repeated here.

[0197] Considering that the storage performance priority of the first storage medium is higher than that of the second storage medium, taking Figure 2 as an example, the first storage medium is the memory and the second storage medium is the hard disk. The read and write speed of the memory is higher than that of the hard disk. If it can be obtained, the training efficiency is higher than that of the hard disk. However, the memory is non-persistent storage. Therefore, the target model parameters may not be obtained and need to be obtained from the hard disk.

[0198] Exemplarily, the stored target model parameter θ is obtained from the memory. If not obtained, the target model parameter θ is obtained from the hard disk.

[0199] The target model parameters are obtained from the first storage medium; if they are not obtained, the target model parameters are obtained from the second storage medium, where the storage performance priority of the first storage medium is higher than that of the second storage medium. This fully utilizes the storage performance priority differences of the storage mediums. In the event of a training anomaly, the target model parameters are preferentially obtained from the storage medium with the higher storage performance priority. At the same time, the target model parameters are ensured to be persistently stored in the storage medium with the higher storage performance priority, which improves the efficiency of model training while ensuring the reliability of model training.

[0200] In an optional embodiment of the present disclosure, the cloud-side device further includes a first communication channel connected to each distributed node and a second communication channel connected to the storage medium;

[0201] In step 504, according to the preset distributed training strategy, multiple distributed nodes are called to perform distributed training on the deep learning model based on the sample data set, including the following specific steps:

[0202] According to a preset distributed training strategy, multiple distributed nodes are called through the first communication channel to perform distributed training on the deep learning model based on the sample data set;

[0203] Correspondingly, in step 504, storing the model parameters of the deep learning model to the storage medium according to the target storage parameters includes the following specific steps:

[0204] According to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the second communication channel.

[0205] The second communication channel connecting the storage medium is a communication channel used for the model parameter storage process, including physical channels and logical channels between distributed nodes and storage media, such as network cards, PCIe topology, optical fibers, read and write channels (data buses) between distributed nodes and storage media, etc.

[0206] The first communication channel connecting to other distributed nodes is used for communication, including physical and logical channels between distributed nodes and storage media. For example, these channels can be network cards, PCIe topologies, optical fibers, high-speed channels between distributed nodes and GPUs, and high-speed channels between distributed nodes. This first communication channel can be implemented using RDMA technology, which improves transmission speed.

[0207] Exemplarily, the second communication channel is a physical channel between the network cards on the 64 distributed nodes and the distributed persistent storage array. According to the target storage model parameter specification P, the model parameters θ of the large language model are stored in the distributed persistent storage array through the second communication channel. The first communication channel is a physical channel between the network cards on the 64 distributed nodes, and a PCIe topology with GPUs inserted on each distributed node, with logical channels for high-speed channels between GPUs and high-speed channels between distributed nodes. Through the first communication channel, the gradient weights on each distributed node are integrated using RDMA technology. Based on the gradient weights on each distributed node, the model parameters θ of the large language model are updated through the gradient update method.

[0208] According to the preset distributed training strategy, multiple distributed nodes are called through the first communication channel to perform distributed training on the deep learning model based on the sample dataset. According to the target storage parameters, the model parameters of the deep learning model are stored to the storage medium through the second communication channel. The two data transmission processes of model parameter storage and communication are isolated to avoid the introduction of additional collective communication and ensure that no interference is caused to distributed training, thereby improving the reliability of distributed training and enhancing the effectiveness of model training.

[0209] Corresponding to the above method embodiments, the present disclosure also provides a deep learning model training system embodiment. FIG6 shows a schematic structural diagram of a deep learning model training system provided by one embodiment of the present disclosure. As shown in FIG6, the system includes a control unit 602 and multiple distributed nodes, the multiple distributed nodes including a first distributed node 604, which is any one of the multiple distributed nodes;

[0210] The control unit 602 is used to obtain an initial deep learning model and a sample data set, construct multiple distributed data based on the deep learning model and the sample data set according to a preset distributed training strategy, and distribute the multiple distributed data to each distributed node;

[0211] The first distributed node 604 is used to perform distributed training on the deep learning model based on the sample data set; and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored therein according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.

[0212] Optionally, the system further includes a persistent storage medium, and the first distributed node 604 includes a non-persistent storage medium;

[0213] Correspondingly, the first distributed node 604 is also used to store the model parameters of the deep learning model to a non-persistent storage medium according to the target storage parameters during the adjustment parameter calculation process of distributed training, so as to transfer the model parameters of the deep learning model from the non-persistent storage medium to the persistent storage medium.

[0214] Optionally, the adjustment parameter is a gradient weight, and the adjustment parameter calculation is a propagation calculation; the first distributed node 604 further includes a first communication channel connected to each distributed node and a second communication channel connected to the storage medium;

[0215] Correspondingly, the first distributed node 604 is further configured to integrate the gradient weights on each distributed node through the first communication channel;

[0216] Correspondingly, the first distributed node 604 is further configured to send the model parameters of the deep learning model from the non-persistent storage medium to the persistent storage medium for storage through the second communication channel.

[0217] In the disclosed embodiment, target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, fully considering the iterative law of distributed training. In the adjustment parameter calculation process of distributed training, the model parameters of the deep learning model are stored according to the target storage parameters. The storage process of the model parameters and the adjustment parameter calculation process in distributed training are overlapped, which fully saves performance overhead. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency.

[0218] The above is a schematic scheme of a deep learning model training system of this embodiment. It should be noted that the technical scheme of this deep learning model training system and the technical scheme of the deep learning model training method described above are based on the same concept. For details not described in detail in the technical scheme of the deep learning model training system, please refer to the description of the technical scheme of the deep learning model training method described above.

[0219] Corresponding to the above method embodiments, the present disclosure also provides an embodiment of a deep learning model training device. FIG7 shows a schematic structural diagram of a deep learning model training device provided by one embodiment of the present disclosure. As shown in FIG7 , the device includes:

[0220] A first acquisition module 702 is configured to acquire an initial deep learning model and a sample dataset;

[0221] The first training module 704 is configured to perform distributed training on the deep learning model based on the sample data set according to a preset distributed training strategy, and store the model parameters of the deep learning model according to the target storage parameters during the calculation of the adjustment parameters of the distributed training, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.

[0222] Optionally, the first training module 704 is further configured to:

[0223] According to the preset distributed training strategy, multiple distributed data are constructed based on the deep learning model and the sample data set, wherein the preset distributed training strategy includes a model parallel training strategy or a data parallel training strategy; the multiple distributed data are distributed to each distributed node; on the first distributed node, propagation calculation is performed based on the distributed data to obtain the gradient weight, wherein the first distributed node is any one of the multiple distributed nodes; based on the gradient weight on each distributed node, the model parameters of the deep learning model are updated, and when the preset training end conditions are met, a trained deep learning model is obtained.

[0224] Optionally, the distributed data includes multiple batches of distributed data;

[0225] The first training module 704 is further configured to:

[0226] On the first distributed node, propagation calculation is performed based on the distributed data of the current batch to obtain the gradient weight;

[0227] Correspondingly, the device further includes:

[0228] The iteration module is configured to update the distributed data of the current batch and return to the step of performing propagation calculation on the first distributed node based on the distributed data of the current batch to obtain the gradient weight.

[0229] Optionally, the device further comprises:

[0230] The storage parameter determination module is configured to predict the number and time overhead of propagation calculations for each batch of distributed data based on the model specification information of the deep learning model and the preset distributed training strategy; based on the number and time overhead of propagation calculations, determine the target storage parameters corresponding to the propagation calculation process.

[0231] Optionally, the device further comprises:

[0232] The gradient weight integration module is configured to integrate the gradient weights on each distributed node through the communication channel between the distributed nodes.

[0233] Optionally, any distributed node includes a first communication channel connected to a storage medium and a second communication channel connected to other distributed nodes;

[0234] The first training module 704 is further configured to:

[0235] Storing the model parameters of the deep learning model to a storage medium via the first communication channel according to the target storage parameters;

[0236] Correspondingly, the gradient weight integration module is further configured as follows:

[0237] The gradient weights on each distributed node are integrated through the second communication channel.

[0238] Optionally, the propagation calculation includes forward propagation calculation and backward propagation calculation;

[0239] The device also includes:

[0240] A forward and reverse storage parameter determination module is configured to determine a first target storage parameter corresponding to a forward propagation calculation process and a second target storage parameter corresponding to a reverse propagation calculation process based on model specification information of the deep learning model and a preset distributed training strategy;

[0241] Correspondingly, the first training module 704 is further configured to:

[0242] During the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameters; during the backward propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameters.

[0243] Optionally, the device further comprises:

[0244] The resumption training module is configured to obtain the stored target model parameters when receiving a resumption training request, wherein the resumption training request is generated after determining that the training anomaly of the distributed training has been recovered, and the target model parameters are the model parameters stored before the training anomaly occurs; based on the target model parameters, the distributed training of the deep learning model is resumed.

[0245] In the disclosed embodiment, target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, fully considering the iterative law of distributed training. In the propagation calculation process of distributed training, the model parameters of the deep learning model are stored according to the target storage parameters. The storage process of the model parameters and the propagation calculation process in the distributed training are overlapped, which fully saves performance overhead. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency.

[0246] The above is a schematic scheme of a deep learning model training device of this embodiment. It should be noted that the technical scheme of the deep learning model training device and the technical scheme of the deep learning model training method described above are based on the same concept. For details not described in detail in the technical scheme of the deep learning model training device, please refer to the description of the technical scheme of the deep learning model training method described above.

[0247] Corresponding to the above method embodiments, the present disclosure also provides an embodiment of a deep learning model training device. FIG8 shows a schematic diagram of the structure of another deep learning model training device provided by one embodiment of the present disclosure. As shown in FIG8, the device is applied to a cloud-side device, which includes multiple distributed nodes and a storage medium; the device includes:

[0248] A second acquisition module 802 is configured to acquire an initial deep learning model and a sample dataset;

[0249] The second training module 804 is configured to call multiple distributed nodes according to a preset distributed training strategy, perform distributed training on the deep learning model based on the sample data set, and store the model parameters of the deep learning model to a storage medium according to a target storage parameter during the distributed training adjustment parameter calculation process, wherein the target storage parameter is determined based on the model specification information of the deep learning model and the preset distributed training strategy;

[0250] The stopping module 806 is configured to trigger the multiple distributed nodes to stop the distributed training of the deep learning model when an abnormality in the deep learning model training is identified, and determine the target model parameters currently stored in the storage medium;

[0251] The parameter acquisition module 808 is configured to acquire the target model parameters from the storage medium when receiving the request to resume training;

[0252] The recovery module 810 is configured to call multiple distributed nodes to perform distributed training on the deep learning model based on target model parameter recovery.

[0253] Optionally, the storage medium includes a plurality of storage media with different storage performances;

[0254] The second training module 804 is further configured to:

[0255] Storing model parameters of the deep learning model in multiple storage media according to target storage parameters and storage performance priorities of the multiple storage media;

[0256] Correspondingly, the parameter acquisition module 808 is further configured to:

[0257] The target model parameters are obtained from the first storage medium; if not obtained, the target model parameters are obtained from the second storage medium, wherein the storage performance priority of the first storage medium is higher than that of the second storage medium.

[0258] Optionally, the cloud-side device further includes a first communication channel connected to each distributed node and a second communication channel connected to the storage medium;

[0259] The second training module 804 is further configured to:

[0260] According to the preset distributed training strategy, multiple distributed nodes are called through the first communication channel to perform distributed training on the deep learning model based on the sample data set; according to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the second communication channel.

[0261] In the embodiment of the present disclosure, target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy, and the iterative law of distributed training is fully considered. In the process of calculating the adjustment parameters of distributed training, the model parameters of the deep learning model are stored in the storage medium according to the target storage parameters, so that the storage process of the model parameters and the adjustment parameter calculation process in distributed training are overlapped, which fully saves performance overhead. The real-time storage of the deep learning model parameters is completed in a manner close to zero performance overhead, so that the deep learning model training has high fault tolerance and high efficiency. When the deep learning model training anomaly is identified, the stored target model parameters are obtained from the storage medium, and the distributed training of the deep learning model is restored, which increases the fault tolerance of the deep learning model training and avoids re-training of the deep learning model. It ensures stability while ensuring training efficiency and reduces training costs.

[0262] The above is a schematic scheme of a deep learning model training device of this embodiment. It should be noted that the technical scheme of the deep learning model training device and the technical scheme of the deep learning model training method described above are based on the same concept. For details not described in detail in the technical scheme of the deep learning model training device, please refer to the description of the technical scheme of the deep learning model training method described above.

[0263] Figure 9 shows a block diagram of a computing device according to an embodiment of the present disclosure. Components of computing device 900 include, but are not limited to, a memory 910 and a processor 920. Processor 920 and memory 910 are connected via a bus 930, and a database 950 is used to store data.

[0264] The computing device 900 also includes an access device 940 that enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC).

[0265] In one embodiment of the present disclosure, the aforementioned components of the computing device 900 and other components not shown in FIG9 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG9 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.

[0266] Computing device 900 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 900 can also be a mobile or stationary server.

[0267] Among them, the processor 920 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned deep learning model training method.

[0268] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-mentioned deep learning model training method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-mentioned deep learning model training method.

[0269] An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned deep learning model training method.

[0270] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-mentioned deep learning model training method are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the above-mentioned deep learning model training method.

[0271] An embodiment of the present disclosure further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned deep learning model training method.

[0272] The above is an illustrative embodiment of a computer program. It should be noted that the technical solution of this computer program and the technical solution of the deep learning model training method described above are based on the same concept. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the deep learning model training method described above.

[0273] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0274] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0275] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present disclosure.

[0276] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0277] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. A deep learning model training method, comprising: Obtain an initial deep learning model and sample dataset; According to the preset distributed training strategy, the deep learning model is distributedly trained based on the sample data set, and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.

2. The method according to claim 1, wherein the adjustment parameter is a gradient weight, and the adjustment parameter calculation is a propagation calculation; The performing distributed training on the deep learning model based on the sample data set according to a preset distributed training strategy includes: According to a preset distributed training strategy, based on the deep learning model and the sample data set, a plurality of distributed data are constructed, wherein the preset distributed training strategy includes a model parallel training strategy or a data parallel training strategy; Distributing the plurality of distributed data to each distributed node; On a first distributed node, performing propagation calculation based on distributed data to obtain a gradient weight, wherein the first distributed node is any one of the multiple distributed nodes; Based on the gradient weights on each distributed node, the model parameters of the deep learning model are updated, and when the preset training end conditions are met, the trained deep learning model is obtained.

3. The method according to claim 2, wherein the distributed data comprises a plurality of batches of distributed data; The step of performing propagation calculation on the first distributed node based on the distributed data to obtain the gradient weight includes: On the first distributed node, a propagation calculation is performed based on the distributed data of the current batch to obtain a gradient weight; After updating the model parameters of the deep learning model based on the gradient weights on each distributed node, the method further includes: The distributed data of the current batch is updated, and the step of performing calculation based on the distributed data of the current batch on the first distributed node to obtain the gradient weight is returned.

4. The method according to claim 3, before performing calculation on the distributed data of the current batch on the first distributed node to obtain the gradient weight, further comprising: For each batch of distributed data, based on the model specification information of the deep learning model and the preset distributed training strategy, predict the number of calculations and time overhead; Based on the number of times and time cost of the calculation, a target storage parameter corresponding to the calculation process is determined.

5. The method according to claim 2, before updating the model parameters of the deep learning model based on the gradient weights on each distributed node, further comprising: The gradient weights on the distributed nodes are integrated through the communication channels between the distributed nodes.

6. The method according to claim 5, wherein any distributed node comprises a first communication channel connected to a storage medium and a second communication channel connected to other distributed nodes; The storing of the model parameters of the deep learning model according to the target storage parameters includes: According to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the first communication channel; The step of integrating the gradient weights on the distributed nodes through the communication channels between the distributed nodes includes: The gradient weights on the distributed nodes are integrated through the second communication channel.

7. The method according to any one of claims 1 to 6, wherein the propagation calculation comprises a forward propagation calculation and a reverse propagation calculation; In the calculation process of the distributed training, before storing the model parameters of the deep learning model according to the target storage parameters, the method further includes: Based on the model specification information of the deep learning model and the preset distributed training strategy, determine a first target storage parameter corresponding to a forward propagation calculation process and a second target storage parameter corresponding to a backward propagation calculation process; The storing of the model parameters of the deep learning model according to the target storage parameters during the calculation process of the distributed training includes: In the forward propagation calculation process, the model parameters of the deep learning model are stored according to the first target storage parameters; During the back propagation calculation process, the model parameters of the deep learning model are stored according to the second target storage parameters.

8. The method according to any one of claims 1 to 7, further comprising: Upon receiving a request to resume training, obtaining stored target model parameters, wherein the request to resume training is generated after determining that a training anomaly of distributed training has been recovered, and the target model parameters are model parameters stored before the training anomaly occurs; Based on the target model parameters, resume distributed training of the deep learning model.

9. A deep learning model training method, applied to a cloud-side device, wherein the cloud-side device includes a plurality of distributed nodes and a storage medium; the method includes: Obtain an initial deep learning model and sample dataset; According to a preset distributed training strategy, calling the multiple distributed nodes, performing distributed training on the deep learning model based on the sample data set, and storing the model parameters of the deep learning model to the storage medium according to the target storage parameters during the adjustment parameter calculation process of the distributed training, wherein the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy; In the case of identifying an abnormality in the deep learning model training, triggering the multiple distributed nodes to stop the distributed training of the deep learning model, and determining the target model parameters currently stored in the storage medium; Upon receiving a request to resume training, obtaining the target model parameters from the storage medium; The multiple distributed nodes are called to perform distributed training on the deep learning model based on the target model parameter recovery.

10. The method according to claim 9, wherein the storage medium comprises a plurality of storage media with different storage performances; The step of storing the model parameters of the deep learning model in the storage medium according to the target storage parameters during the calculation process of the distributed training includes: storing the model parameters of the deep learning model in the multiple storage media according to the target storage parameters and the storage performance priorities of the multiple storage media; The acquiring the target model parameters from the storage medium includes: Acquiring the target model parameters from a first storage medium; If not obtained, the target model parameters are obtained from a second storage medium, wherein the storage performance priority of the first storage medium is higher than that of the second storage medium.

11. The method according to claim 9, wherein the cloud-side device further comprises a first communication channel connected to each distributed node and a second communication channel connected to the storage medium; According to a preset distributed training strategy, calling the multiple distributed nodes, and performing distributed training on the deep learning model based on the sample data set, including: According to a preset distributed training strategy, calling the plurality of distributed nodes through the first communication channel, and performing distributed training on the deep learning model based on the sample data set; The storing the model parameters of the deep learning model to the storage medium according to the target storage parameters includes: According to the target storage parameters, the model parameters of the deep learning model are stored in the storage medium through the second communication channel.

12. A deep learning model training system, the system comprising a management and control unit and a plurality of distributed nodes, the plurality of distributed nodes comprising a first distributed node, the first distributed node being any one of the plurality of distributed nodes; The control unit is used to obtain an initial deep learning model and a sample data set, construct a plurality of distributed data based on the deep learning model and the sample data set according to a preset distributed training strategy, and distribute the plurality of distributed data to each distributed node; The first distributed node is used to perform distributed training on the deep learning model based on the sample data set; and in the process of calculating the adjustment parameters of the distributed training, the model parameters of the deep learning model are stored according to the target storage parameters. Among them, the target storage parameters are determined based on the model specification information of the deep learning model and the preset distributed training strategy.

13. The system of claim 12, wherein the system further comprises a persistent storage medium, and the first distributed node comprises a non-persistent storage medium; The first distributed node is also used to store the model parameters of the deep learning model to the non-persistent storage medium according to the target storage parameters during the adjustment parameter calculation process of the distributed training, so as to transfer the model parameters of the deep learning model from the non-persistent storage medium to the persistent storage medium.

14. The system according to claim 13, wherein the adjustment parameter is a gradient weight, and the adjustment parameter calculation is a propagation calculation; The first distributed node also includes a first communication channel connected to each distributed node and a second communication channel connected to the storage medium; The first distributed node is further used to integrate the gradient weights on each distributed node through the first communication channel; The first distributed node is also used to send the model parameters of the deep learning model from the non-persistent storage medium to the persistent storage medium for storage through the second communication channel.

15. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 11 are implemented.

16. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.

17. A computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Deep learning model training method and deep learning model training system

    CN117669700B

  • Method and device for neural network machine learning model training

    CN109754060A

  • Deep learning model training method and system

    CN111788585A

  • Distributed training method for large-scale deep neural network

    CN113515370A

  • Training device and method of neural network model and related equipment

    CN113705801A

Cited By

  • Large model distributed parallel training method, program product, computing node and storage medium

    CN121412032A

  • Data communication method and device under distributed training task, equipment, storage medium and program product

    CN121792535A

  • Partition classification parallel training method for large mountain torrent risk prediction model

    CN122112861A

  • Training method and system, parameter updating method, electronic equipment and storage medium

    CN122334383A

  • AI large model training method based on deep learning

    CN122389943A