Distributed training method and device of neural network model, medium and cluster

By detecting the model scale, performing quantitative processing and distributed flow parallel training, the problems of low computational efficiency and resource utilization in deep learning model training are solved, and efficient and stable model training is achieved.

CN120087438APending Publication Date: 2025-06-03KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510157281.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

With the increase in the number of parameters of deep learning models, existing distributed training methods are difficult to effectively improve the computing efficiency and resource utilization. At the same time, low-precision quantization leads to accuracy loss, affecting the convergence performance of the model.

Method used

A distributed training method for neural network models is proposed. By detecting whether the model scale of the model to be trained exceeds the processing scale of the distributed cluster, if it exceeds the processing scale, quantization processing is performed and the quantization model is split into multiple model partitions, and distributed parallel training is respectively arranged in multiple flow parallel groups. During the reverse calculation process, each flow group independently maintains the gradient scaling value to update the model weights.

Benefits of technology

Through quantitative processing and distributed flow parallel training, the efficiency and resource utilization of model training are improved, the accuracy loss caused by low-precision quantization is avoided, and the convergence performance of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087438A_ABST
    Figure CN120087438A_ABST
Patent Text Reader

Abstract

The invention provides a distributed training method and device of a neural network model, a medium and a cluster, and relates to the field of artificial intelligence, in particular to the field of chips. According to the specific implementation scheme, when it is detected that the model scale of a to-be-trained model is larger than the processing scale of a distributed cluster, the model is quantized to obtain a quantized model; splitting the quantitative model into a plurality of model partitions and configuring the model partitions in each pipeline parallel group; when forward calculation is carried out, a training error value is generated through common cooperation among the pipeline parallel groups; and in the process of carrying out reverse calculation on the quantitative model according to the training error value, the model weight is updated through each pipeline parallel group according to the gradient scaling value independently maintained in the group. The parallel groups are used as granularity, each parallel group maintains the respective gradient scaling value and updates the weight in the respective model partition by using the gradient scaling value, so that the model precision and the calculation overhead can be effectively balanced, and the efficiency, the precision and the reliability of distributed training are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to the fields of artificial intelligence, chips, etc. Specifically, it relates to a distributed training method for a neural network model, a distributed training device for a neural network model, a device, a non-transitory computer-readable storage medium, a computer program product, and a distributed cluster. Background Art

[0002] With the wide application of deep learning and large models, the number of model parameters continues to grow, and the computing time and resources required for deep learning training also continue to increase. To improve computing efficiency and save resources, low-precision quantization techniques are generally adopted. However, low-precision quantization will lead to accuracy loss, which in turn affects the convergence performance of the model. Therefore, error quantization techniques are usually further adopted to effectively improve the training accuracy. Summary of the Invention

[0003] The present disclosure provides a distributed training method for a neural network model, a distributed training device for a neural network model, a device, a non-transitory computer-readable storage medium, a computer program product, and a distributed cluster.

[0004] According to one aspect of the present disclosure, there is provided a distributed training method for a neural network model, which is executed by a distributed cluster. The distributed cluster includes multiple pipelining parallel groups, and each pipelining parallel group includes multiple devices. The method includes:

[0005] When it is detected that the model scale of the model to be trained is greater than the processing scale of the distributed cluster, perform quantization processing on the model to be trained to obtain a quantized model;

[0006] Split the quantized model into multiple model partitions, and respectively configure each model partition in each pipelining parallel group to perform distributed pipelining parallel training for the quantized model;

[0007] During the forward calculation of the quantized model according to the quantized training samples, generate a training error value through the joint cooperation between each pipelining parallel group;

[0008] During the backward calculation of the quantized model according to the training error value, update the model weights in the configured model partitions by each pipelining parallel group according to the gradient scaling values independently maintained within the group.

[0009] According to another aspect of the present disclosure, there is provided a device for executing the distributed training method for a neural network model according to any one of the embodiments of the present disclosure.

[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a device to execute the distributed training method of the neural network model according to any one of the embodiments of the present disclosure.

[0011] According to another aspect of the present disclosure, there is provided a distributed cluster including a plurality of devices, and each device is used to execute the distributed training method of the neural network model according to any one of the embodiments of the present disclosure.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0013] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0014] Figure 1 is a schematic diagram of a distributed training method of a neural network model according to an embodiment of the present disclosure;

[0015] Figure 2 is a schematic diagram of another distributed training method of a neural network model according to an embodiment of the present disclosure;

[0016] Figure 3 is a schematic diagram of yet another distributed training method of a neural network model according to an embodiment of the present disclosure;

[0017] Figure 4 is a schematic diagram of a model training method using a unified scaling value provided by the related art;

[0018] Figure 5 is a schematic diagram of the overall execution process in the distributed training of a neural network model applicable to the embodiments of the present disclosure;

[0019] Figure 6 is a schematic diagram of the parallel group execution process in the distributed training of a neural network model applicable to the embodiments of the present disclosure;

[0020] Figure 7 is a schematic diagram of a distributed training device of a neural network model according to an embodiment of the present disclosure;

[0021] Figure 8 is a schematic diagram of a device according to an embodiment of the present disclosure;

[0022] Figure 9 is a schematic diagram of a distributed cluster according to an embodiment of the present disclosure. Detailed Embodiments

[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0024] In a distributed training scenario, multiple machines collaborate to complete the computational tasks of multiple consecutive layers in a neural network through pipelining parallelism, enabling the construction of an efficient distributed computing process. However, as the number of model parameters continues to grow, the pipelining parallelism approach is still insufficient to meet the efficiency requirements of distributed training. Among the three main steps of model training: forward computation, gradient backpropagation, and parameter update, by using low-precision training, that is, by using a lower-precision data type, such as FP16 (Floating Point 16, 16-bit floating-point number), memory occupancy can be reduced and computation can be accelerated. However, this approach will result in a certain loss of precision and reduce the convergence performance of the model during distributed training.

[0025] In related technologies, in order to balance precision and computational efficiency during model training, error quantization techniques are usually adopted. Specifically, this method introduces a fixed gradient scaling value during the backpropagation process, calculates the gradient by magnifying the error value, and then restores it to its original size. However, this method is not applicable to all operators or computational layers with different data distributions. When facing operators or computational layers with significant differences in data distribution, the fixed gradient scaling value is prone to failure, thus unable to effectively maintain model precision and the convergence of the model. In addition, implementing corresponding error quantization schemes for each operator separately not only increases the difficulty of engineering implementation but also significantly improves the complexity of the entire system.

[0026] Figure 1 is a schematic diagram of a distributed training method for a neural network model according to an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the case of parallel group quantization in the distributed training of a neural network model. This method can be executed by a distributed training device of the neural network model. The device can be implemented in a hardware manner and is generally configured in a distributed cluster including multiple pipelining parallel groups, and each pipelining parallel group includes multiple devices.

[0027] Correspondingly, as Figure 1 shown, the method may specifically include:

[0028] S110. When it is detected that the model scale of the model to be trained is greater than the processing scale of the distributed cluster, perform quantization processing on the model to be trained to obtain a quantized model.

[0029] Generally speaking, in a distributed training scenario, the model scale can be specifically understood as: the storage space and computing resources required for the number of model parameters, activation values, and intermediate calculation results, etc. For large neural network models, the number of parameters may reach billions or even trillions, making it difficult for a single computing node to accommodate and process. The processing scale can be specifically understood as: the total storage capacity and computing power of all computing nodes in the distributed cluster. When the model scale exceeds the processing scale of the cluster, it means that the existing resources of the cluster alone cannot directly complete the training of the model.

[0030] Quantization processing can be specifically understood as: the process of converting model parameters and activation values from high-precision numerical values (such as 32-bit floating-point numbers FP32) to low-precision numerical values (such as FP16), in order to reduce the memory occupancy of the model, making the model easier to be split into multiple computing nodes and improving the computing efficiency. Specifically, it can be: according to the distribution range of model parameters or activation values, through quantization methods such as symmetric quantization or asymmetric quantization, determine a suitable quantization interval, map the floating-point numerical values to low-precision discrete values, and obtain a quantized model.

[0031] Among them, the model to be trained can be applied in the fields of computer vision, natural language processing, medical treatment, or autonomous driving;

[0032] The field of computer vision includes image classification, object detection, or face recognition; the field of natural language processing includes text classification, machine translation, or speech recognition, etc.

[0033] S120. Split the quantized model into multiple model partitions, and configure each model partition in each pipelined parallel group respectively to perform distributed pipelined parallel training for the quantized model.

[0034] Generally speaking, in a distributed training scenario, the model partition can be specifically understood as: splitting the quantized model into multiple parts, and each partition contains a part of the layers of the model. The pipelined parallel group can be specifically understood as: consisting of multiple parallel groups, and there is a sequential order of data calculation and communication between each parallel group, and each parallel group consists of multiple devices. By configuring each model partition into different pipelined parallel groups respectively, each group only processes one partition of the model. During the training process, data passes through each parallel group in sequence in the pipeline, and each group sequentially completes the calculation of the layers it is responsible for. Among them, during forward propagation, data passes through each partition in sequence from the input end; during backward propagation, the gradient propagates back from the output end to the input end in reverse.

[0035] Specifically, the total number of layers of each neural network layer in the model can be counted. For example, assume the model has 12 layers. According to the hardware resources and requirements of distributed training, a basic partition number is preset. For example, assume the preset basic partition number is 4. Therefore, the total number of model layers can be divided by the preset basic partition number to obtain the average number of layers contained in each partition. For example, the average number of layers is 3 layers / partition. In addition, the computational complexity of each layer (such as the number of parameters and the amount of computation, etc.) can also be counted. The partitions with multiple layers with large computational amounts or high memory occupancy rates in the same partition are further split, and the partitions with layers with small computational amounts or low memory occupancy rates are merged, so that the computational amounts or memory occupancy rates between different partitions are similar, and the partition number and the number of layers in each partition are adjusted flexibly according to the actual situation. Each partition is assigned to different pipeline parallel groups. Each parallel group is responsible for processing a partition of the model. In a distributed environment, training is executed. Through pipeline parallelism, data sequentially passes through each parallel group in the pipeline, and each group sequentially completes the calculation of the layers it is responsible for.

[0036] S130. During the forward calculation of the quantization model based on the quantization training samples, through the joint cooperation between the pipeline parallel groups, a training error value is generated.

[0037] Specifically, during the forward calculation process, the input data, through the joint cooperation between the pipeline parallel groups, sequentially passes through the model partitions corresponding to each pipeline parallel group. Each partition completes the calculation of the layers it is responsible for and passes the output result to the next parallel group. Finally, after the last parallel group completes the calculation of the layers it is responsible for, the output of the model is generated. The training error value is calculated using the output of the model and the target value through a loss function (such as cross-entropy loss or mean squared error, etc.). Among them, the calculated training error value will be used for backpropagation to update the weights of the model.

[0038] S140. During the backward calculation of the quantization model based on the training error value, each pipeline parallel group updates the model weights in the configured model partition according to the gradient scaling value independently maintained within the group.

[0039] Generally speaking, in deep learning, the gradient scaling value is a coefficient used to adjust the magnitude of the gradient. It is used to prevent the vanishing gradient problem where, during the backpropagation process in a deep neural network, as the number of network layers increases, the gradient values gradually decrease layer by layer and eventually approach zero, resulting in very slow or even almost stagnant weight updates near the input layer. It also prevents the gradient explosion problem where the gradient values continuously increase during multi-layer transmission, leading to overly drastic weight updates in the network, even numerical overflow, causing unstable updates of model parameters and non-convergence of the training process. When training using a low-precision data type (such as FP16), due to the smaller effective data range of FP16 compared to FP32, problems of overflow or underflow are likely to occur. Additionally, at low precision, smaller gradients may be discarded as zero due to insufficient precision, resulting in gradient underflow, and corresponding overflow problems also occur when the numerical value exceeds the representable range of half-precision.

[0040] Specifically, during the backpropagation process, the training error value starts from the last parallel group and is propagated backward to the previous parallel groups in sequence. Each parallel group calculates the gradient of the layers it is responsible for, multiplies the calculated gradient by the gradient scaling value of that parallel group to adjust the magnitude of the gradient, and uses the scaled gradient to update the weights in the model partition that the parallel group is responsible for. Among them, when calculating using FP16, due to the lower precision, the calculated loss value often tends to be smaller. To ensure that the gradient has a sufficient numerical range during the backpropagation process and avoid overly small gradients that cannot effectively update the weights, it is necessary to amplify the loss value through gradient scaling. Therefore, the initial value of the gradient scaling value is often set to a number greater than 1, which can amplify the loss value and thus obtain a larger gradient value during backpropagation. Optionally, it can be set to the 20th power of 2, which is a fixed value. Additionally, in practical applications, the gradient scaling value can be dynamically adjusted according to the numerical stability during the training process. For example, if gradient overflow is detected (i.e., the gradient value is too large, resulting in numerical instability), the scaling value is decreased; if the gradient is too small, the scaling value is increased.

[0041] Based on this, the technical solution of the embodiments of the present disclosure obtains a quantized model by performing model quantization when it is detected that the model scale of the model to be trained is larger than the processing scale of the distributed cluster; splits the quantized model into multiple model partitions and configures them in each pipelined parallel group; during forward calculation, through the joint cooperation between each pipelined parallel group, a training error value is generated; during the process of performing backward calculation on the quantized model according to the training error value, each pipelined parallel group updates the model weights according to the gradient scaling value independently maintained within the group. Taking the parallel group as the granularity, each parallel group can better adapt to the numerical distribution of different partitions by maintaining its own gradient scaling value and using it to update the weights in its own model partition, avoiding the accuracy loss caused by the globally unified scaling value, improving the stability and convergence of model training, and effectively achieving a balance between model accuracy and computational overhead. The method based on the parallel group makes the implementation of distributed training more modular, simplifies the complexity of engineering implementation, and improves the efficiency, accuracy, and reliability of distributed training.

[0042] In an optional implementation manner of this embodiment, detecting that the model scale of the model to be trained is larger than the processing scale of the distributed cluster may include:

[0043] According to the model scale of the model to be trained, predict the target storage requirement during training of the model to be trained;

[0044] Obtain the standard storage size that the distributed cluster can provide;

[0045] If the standard storage size does not meet the target storage requirement, it is determined that the model scale of the model to be trained is larger than the processing scale of the distributed cluster.

[0046] Generally speaking, the target storage requirement can be specifically understood as: predicting the total storage space required during model training according to the number of model parameters, activation values, and intermediate calculation results, etc., which can be specifically completed by analyzing the model structure and training configuration. For example, assuming the model has 1 billion parameters and each parameter occupies 2 bytes (FP16), then the storage space required for the model parameters is 10 9 ×2 bytes = 2 GB.

[0047] The standard storage size can be specifically understood as: the available storage capacity of each computing node in the distributed cluster, which can be specifically obtained by querying the cluster configuration or performing a runtime check. For example, assuming the cluster has 10 computing nodes and each computing node has 16 GB of video memory, then the total video memory of the cluster is 160 GB.

[0048] It can be understood that by comparing the target storage requirement with the standard storage size, if the target storage requirement exceeds the standard storage size of the cluster, it is determined that the model scale is larger than the processing scale of the cluster, and the current distributed cluster is not sufficient to support the training of the current neural network model. It is necessary to perform quantization processing using the distributed training method of the neural network model proposed in this embodiment of the disclosure and perform distributed training in the form of parallel groups.

[0049] By detecting the model scale of the model to be trained, predicting the target storage requirement during model training, and simultaneously obtaining the standard storage size that the distributed cluster can provide, it is possible to effectively determine whether the model scale exceeds the processing capacity of the cluster. If the standard storage size does not meet the target storage requirement, it is determined that the model scale is larger than the processing scale of the cluster. This method can identify potential resource bottlenecks in advance, thereby providing a basis for subsequent model quantization and distributed training, ensuring the smooth progress of the training process, improving resource utilization efficiency, avoiding training failures caused by insufficient resources, and improving the reliability of distributed training.

[0050] Figure 2 It is a schematic diagram of another distributed training method of a neural network model provided according to an embodiment of the disclosure. This embodiment is a refinement of the operation of "splitting the quantization model into multiple model partitions" in the above embodiment. Specifically, it can be: according to the total number of neural network layers included in the quantization model and a preset basic partition number, determine the number of neural network layers included in each model partition; split the quantization model into multiple model partitions according to the number of neural network layers.

[0051] Correspondingly, as Figure 2 shown, the method may specifically include:

[0052] S210. When it is detected that the model scale of the model to be trained is larger than the processing scale of the distributed cluster, perform quantization processing on the model to be trained to obtain a quantization model.

[0053] S220. According to the total number of neural network layers included in the quantization model and a preset basic partition number, determine the number of neural network layers included in each model partition.

[0054] Specifically, count the total number of neural network layers in the model. For example, assume the model has 12 layers. According to the hardware resources and requirements of distributed training, preset a basic partition number. For example, assume the preset basic partition number is 4. Divide the total number of model layers by the preset basic partition number to obtain the average number of layers included in each partition, such as 3 layers / partition.

[0055] S230. Split the quantization model into multiple model partitions according to the number of neural network layers.

[0056] Specifically, in the order of the layers of the neural network model, each layer of the quantization model is successively split into multiple model partitions according to the number of neural network layers included in each calculated model partition. For example, Partition 1: includes the 1st - 3rd layers of the quantization model; Partition 2: includes the 4th - 6th layers of the quantization model; Partition 3: includes the 7th - 9th layers of the quantization model; Partition 4: includes the 10th - 12th layers of the quantization model.

[0057] S240. Configure each of the model partitions in each stream parallel group respectively to perform distributed stream parallel training for the quantization model.

[0058] S250. During the forward calculation of the quantization model based on the quantization training samples, generate a training error value through the joint cooperation between each of the stream parallel groups.

[0059] S260. During the backward calculation of the quantization model based on the training error value, update the model weights in the configured model partitions by each of the stream parallel groups according to the gradient scaling values independently maintained within the group.

[0060] Based on this, the technical solution of the embodiments of the present disclosure determines the number of neural network layers included in each model partition according to the total number of neural network layers included in the quantization model and a preset basic partition number; splits the quantization model into multiple model partitions according to the number of neural network layers, realizing efficient splitting and management of the model, ensuring load balancing of the model in distributed training, avoiding performance bottlenecks caused by over - heavy tasks on some devices, improving resource utilization rate, enabling computing resources to be more reasonably allocated to each model partition, facilitating subsequent pipeline parallel computing, and further improving the efficiency, accuracy, and reliability of distributed training.

[0061] In an optional implementation manner of this embodiment, splitting the quantization model into multiple model partitions according to the number of neural network layers specifically includes:

[0062] Pre - split the quantization model into multiple alternative partitions according to the number of neural network layers;

[0063] Starting from the pre - split alternative partitions, perform at least one pre - test on the quantization model by combining the use of quantization training samples, and perform at least one fine - tuning process on each of the alternative partitions according to the pre - test results to obtain the multiple model partitions that match the quantization model.

[0064] Specifically, according to the total number of layers of the model and the preset number of basic partitions, the quantized model is pre-split into multiple alternative partitions. For example, assume the model has 12 layers and the preset number of basic partitions is 4, then each alternative partition contains 3 layers. Starting from the pre-split alternative partitions, pre-tests are performed on each alternative partition using quantized training samples, and the performance metrics of each partition (such as computing time or memory usage, etc.) are recorded to evaluate the performance of each alternative partition, including computing load, memory occupancy, and communication overhead, etc. According to the performance evaluation results, fine-tuning is performed on the alternative partitions. For example, if the computing load of a certain partition is too heavy, its number of layers can be appropriately reduced; if the computing load of a certain partition is too light, its number of layers can be appropriately increased. During the fine-tuning process, it can be specified that each partition contains at least 2 layers, and the number of parallel groups cannot be less than the preset number of basic partitions. Through at least one fine-tuning process, multiple model partitions matching the quantized model are finally obtained. These partitions will be configured on different computing nodes for distributed training.

[0065] By pre-splitting the model according to the number of neural network layers and combining fine-tuning processing based on the pre-test results, the model partitions can be optimized more precisely, ensuring the balance of computing load and memory occupancy of each partition, improving resource utilization, and reducing communication overhead during the training process. Through pre-tests and fine-tuning, the number of layers of the partitions can be dynamically adjusted to ensure the optimal performance of each partition, thus significantly improving the efficiency, reliability of distributed training, and the final performance of the model.

[0066] In an alternative implementation manner of this embodiment, starting from the pre-split alternative partitions, at least one pre-test is performed on the quantized model by combining the use of quantized training samples, and according to the pre-test results, at least one fine-tuning process is performed on each of the alternative partitions to obtain the multiple model partitions matching the quantized model, which may include:

[0067] Pre-configure each of the alternative partitions in each of the pipelined parallel groups;

[0068] Input at least one quantized training sample into the quantized model that has been pre-configured for pre-testing, and obtain the numerical distribution of each neural network layer in each pipelined parallel group;

[0069] If it is determined that the numerical distribution of each neural network layer in at least one target pipelined parallel group does not meet the intra-group distribution consistency condition, then fine-tune each of the alternative partitions according to the numerical distribution;

[0070] Return to execute the operation of pre-configuring each of the alternative partitions in each of the pipelined parallel groups until the end iteration condition is met;

[0071] Each of the alternative partitions obtained when the iteration ends is determined as each of the model partitions matching the quantization model.

[0072] Specifically, the pre-split alternative partitions are respectively configured into each pipeline parallel group, and each parallel group is responsible for processing one alternative partition. At least one quantization training sample is input into the pre-configured quantization model for pre-testing, and the numerical distributions of each neural network layer in each pipeline parallel group are obtained. The numerical distribution may include statistical information such as mean, variance, and gradient magnitude. Check whether the numerical distributions in each pipeline parallel group satisfy the intra-group distribution consistency condition, such as whether the mean differences, variance differences, and gradient magnitude differences of each layer are within the corresponding preset ranges. If it is found that the numerical distributions in some parallel groups do not satisfy the consistency condition, fine-tuning is required. According to the results of the numerical distribution, fine-tune the alternative partitions that do not satisfy the consistency condition, move some layers from one partition to another to increase or decrease the number of layers in the partition where the numerical distribution does not satisfy the consistency condition, thereby optimizing the numerical distribution. For example: If the variance of the numerical distribution of a certain partition is large, it means that the computational load of this partition is heavy, and some layers can be moved to other partitions to reduce the burden on this partition; if the variance of the numerical distribution of a certain partition is small, it means that the computational load of this partition is light, and some layers can be moved to other partitions to balance the computational load. That is: If the variance of the numerical distribution of a certain partition is large, some layers can be moved to the partition with a small variance of the numerical distribution. If the variance of the numerical distribution of a certain partition is small, some layers can be moved to the partition with a large variance of the numerical distribution.

[0073] Re-execute the operations of pre-configuration, pre-testing, and fine-tuning until the end iteration condition is met. The end iteration condition may include: the numerical distributions of all parallel groups satisfy the consistency condition, or even if the numerical distributions of some parallel groups still do not fully satisfy the consistency condition, but the preset number of iterations has been reached, or the change in the numerical distribution is less than a certain threshold, and it is considered to have converged. The alternative partitions obtained when the iteration ends are determined as the final model partitions matching the quantization model.

[0074] Through the iterative optimization process of pre-configuration, pre-testing, and fine-tuning, the model partitions can be dynamically adjusted to ensure that the numerical distribution of each partition satisfies the consistency condition, which not only improves the efficiency of distributed training, but also optimizes the resource utilization rate and reduces the communication overhead during the training process. Through iterative optimization, the partitions can be gradually adjusted to finally obtain the optimal partition scheme matching the quantization model, thereby significantly improving the training efficiency, accuracy, and reliability of the model.

[0075] In an alternative implementation manner of this embodiment, fine-tuning each of the alternative partitions according to the numerical distribution specifically includes:

[0076] Obtain the current numerical distributions of the current neural network layers in the current target pipeline parallel group;

[0077] According to each of the current numerical distributions, split each of the current neural network layers into multiple groups that satisfy the distribution consistency condition; wherein, each group contains at least one current neural network layer;

[0078] According to the positions of each of the groups in the quantization model and the number values of the neural network layers included in each of the groups, perform a merging process on each of the groups and adjacent alternative partitions;

[0079] If the number of remaining groups after the merging process is multiple, determine each of the multiple groups as different alternative partitions.

[0080] Specifically, obtain the current numerical distributions (such as mean, variance, and gradient magnitude, etc.) of the neural network layers in the current target pipeline parallel group, split the neural network layers into multiple groups, and the layers in each group satisfy the distribution consistency condition (such as the mean difference, variance difference, and gradient magnitude difference of each layer are within a preset range). According to the positions of each group in the quantization model and the number of neural network layers included in the group, perform a merging process on the group and adjacent alternative partitions to optimize the partitions and ensure that the numerical distribution of each partition is more balanced. Specifically, it can be: if the numerical distribution gaps between adjacent groups are not large, and the number of neural network layers included after the group merging is similar to the number of neural network layers included in other groups, these groups can be merged into one partition. For example: Group 1 includes Layer 1 and Layer 2, Group 2 includes Layer 3. If the numerical distribution gaps between Group 1 and Group 2 are not large, and the number of layers is similar, they can be merged into a new Partition 1.

[0081] If the numerical distribution gaps between adjacent groups are large, or the number of neural network layers included in the groups is quite different, no merging process is performed, but the groups are kept independent. For example: Group 1 includes Layer 1 and Layer 2, Group 2 includes Layer 3. If the numerical distribution gaps between Group 1 and Group 2 are large, or the number of layers is quite different, they are respectively used as independent Partition 1 and Partition 2.

[0082] Optionally, if after splitting, there are groups in the current grouping method where the numerical distribution does not meet the within-group consistency condition, the following method can be used for merging processing. Generally speaking, the numerical distributions of adjacent neural network layers are generally similar. Due to the equal distribution method, neural network layers that originally met the within-group distribution consistency condition may be divided into different parallel groups, resulting in non-compliance with the within-group distribution consistency condition. Therefore, starting from the neural network layer that does not meet the within-group distribution consistency condition, it can be merged with adjacent neural network layers with similar forward or backward numerical distributions. If after merging, the number of neural networks in the parallel group is greater than the preset threshold, no merging processing is performed. Instead, the elements in the original parallel group that do not meet the within-group distribution consistency condition are split so that the numerical distributions of the split parallel groups meet the within-group distribution consistency condition and the number of network layers contained in each group is within a reasonable range.

[0083] After completing the merging process, the final grouping is determined as multiple model partitions that match the quantization model. If the number of remaining groups after completing the merging process is multiple, these groups are respectively determined as different alternative partitions.

[0084] In the merging process, by considering the number of groups and the consistency of the numerical distribution, the alternative partitions are finely adjusted to avoid partition imbalance caused by excessive quantity differences. The model partitions can be dynamically optimized to ensure that the numerical distribution of each partition meets the consistency condition, and the number of layers contained in each partition is similar, improving the efficiency of distributed training, optimizing resource utilization, and reducing the communication overhead during the training process. By dynamically adjusting the number of layers and moving layers of the partitions, the numerical distribution can be gradually optimized, and finally the optimal partition scheme that matches the quantization model can be obtained, thus significantly improving the training performance and final accuracy of the model.

[0085] Figure 3 It is a schematic diagram of another method for distributed training of a neural network model provided according to an embodiment of the present disclosure. This embodiment refines the operation of "during the reverse calculation of the quantization model according to the training error value, each of the pipelined parallel groups updates the model weights in the configured model partition according to the gradient scaling value independently maintained within the group" in the above embodiment. Specifically, if the current pipelined parallel group is the first parallel group in the reverse propagation direction, perform target scaling calculation on the training error value according to the gradient scaling value within the group to obtain the current update basis value; if the current pipelined parallel group is not the first parallel group, perform target scaling calculation on the gradient output value received from the previous pipelined parallel group in the reverse propagation direction according to the gradient scaling value within the group to obtain the current update basis value; update the model weights in the configured model partition according to the current update basis value.

[0086] Correspondingly, as Figure 3As shown, the method may specifically include:

[0087] S310. When it is detected that the model scale of the model to be trained is greater than the processing scale of the distributed cluster, perform quantization processing on the model to be trained to obtain a quantized model.

[0088] S320. Split the quantized model into multiple model partitions, and respectively configure each model partition in each pipelined parallel group to perform distributed pipelined parallel training for the quantized model.

[0089] S330. In the process of forward calculation of the quantized model according to the quantized training samples, generate a training error value through the joint cooperation between each pipelined parallel group.

[0090] S340. If the current pipelined parallel group is the first parallel group in the reverse propagation direction, perform target scaling calculation on the training error value according to the gradient scaling value within the group to obtain the current update basis value.

[0091] Generally speaking, in the reverse propagation process, the first parallel group can be specifically understood as: the group that first receives gradient information, is responsible for processing the last part of the model layers, and passes the gradient information to the next parallel group. Since the first parallel group is the first group to process gradients, it is necessary to perform preliminary processing on the training error value to ensure that the gradient information will not become invalid due to numerical problems during subsequent transmission.

[0092] S350. If the current pipelined parallel group is not the first parallel group, perform target scaling calculation on the gradient output value received from the previous pipelined parallel group in the reverse propagation direction according to the gradient scaling value within the group to obtain the current update basis value.

[0093] Generally speaking, in the reverse propagation process, the non - first parallel group can be specifically understood as: the group that is not the first to receive gradient information, needs to process the gradient information received from the previous parallel group, ensure that these gradient information will not become invalid due to numerical problems (such as gradient vanishing or gradient explosion) during transmission, and continue to pass it backward.

[0094] S360. Update the model weights in the configured model partitions according to the current update basis value.

[0095] Specifically, multiply the training error value by the gradient scaling value within the first parallel group to obtain the scaled error value, which is the update basis value for the first parallel group, and then perform subsequent gradient calculations to obtain the gradient output value and the gradient update value. Among them, the gradient update value refers to the gradient value calculated for each layer within each parallel group during the backpropagation process and is used to update the weights of that layer. The gradient output value refers to the gradient value calculated for the last layer of each parallel group during the backpropagation process and is used to be passed to the next parallel group.

[0096] Divide the gradient update value by the gradient scaling value within the first parallel group for restoration, and update the model weights in the configured model partition through the optimizer. Divide the gradient output value by the gradient scaling value within the first parallel group for restoration, and pass the restored gradient output value to the next parallel group.

[0097] For non-first parallel groups, scale the gradient output value received from the previous parallel group by multiplying it by the gradient scaling value within the current parallel group to obtain the current update basis value, and then perform subsequent gradient calculations to obtain the gradient output value and the gradient update value. Divide the gradient update value by the gradient scaling value within the current parallel group for restoration, and update the model weights in the configured model partition through the optimizer. Divide the gradient output value by the gradient scaling value within the current parallel group for restoration, and pass the restored gradient output value to the next parallel group. Repeat the weight update and the operation of passing the restored gradient output value for non-first parallel groups until the last parallel group completes the weight update.

[0098] Based on this, the technical solution of the embodiments of the present disclosure provides corresponding weight update strategies for each pipelined parallel group by introducing a gradient scaling mechanism during the backpropagation process, thereby significantly improving the accuracy and stability of distributed training. For the first parallel group in the backpropagation direction, scale the training error value based on the gradient scaling value within the group to obtain the basis value for update; for other parallel groups, perform scaling calculations based on the gradient output value passed from the previous group. This update strategy can enhance the adaptability of distributed training to different data distributions, avoid accuracy loss caused by a unified scaling value, and ensure that each parallel group can independently adjust its weights according to its own situation, further optimizing the model performance. In addition, this method simplifies the synchronization operation in distributed training at the granularity of parallel groups, reduces the communication overhead, and the method based on parallel groups makes the implementation of distributed training more modular, simplifies the complexity of engineering implementation, and is especially suitable for the distributed training scenario of neural network models, improving the efficiency, accuracy, and reliability of distributed training.

[0099] In an alternative implementation of this embodiment, updating the model weights in the configured model partition according to the current update basis value may include:

[0100] Calculating a gradient update value and a gradient output value that match the current pipeline parallel group according to the current update basis value;

[0101] Determining whether the gradient update value is within a preset reasonable data range;

[0102] If so, performing a gradient reduction calculation on the gradient update value and the gradient output value according to the gradient scaling value within the group;

[0103] Updating the model weights in the configured model partition of the current pipeline parallel group according to the reduced gradient update value;

[0104] Outputting the reduced gradient output value to the next pipeline parallel group in the backpropagation direction of the current pipeline parallel group.

[0105] Specifically, before performing the gradient reduction calculation, it is necessary to determine whether the gradient update value is within a preset reasonable data range to avoid problems such as unstable numerical values and gradient disappearance or explosion caused by too large or too small gradients. Specifically, it can be: according to the current update basis value (i.e., the scaled gradient), calculate the gradient update value ▽W i ’ that matches the current pipeline parallel group, and calculate the gradient output value O i ’ that will be passed to the next parallel group. Among them, there are multiple layers within each parallel group, and each layer has its own gradient update value. Each gradient update value needs to be checked for overflow to ensure numerical stability. Each parallel group has only one gradient output value, corresponding to the last layer of the parallel group.

[0106] Check whether the gradient update value ▽W i ’ is within a reasonable range (i.e., whether it exceeds a certain preset threshold and numerical overflow occurs). If the gradient update value is within a reasonable range, perform a reduction calculation on the gradient update value and the gradient output value according to the gradient scaling value within the group, that is, ▽W i = ▽W i ’ / S i and O i = O i ’ / S i i i where S i is the gradient scaling value within the current parallel group, and ▽W iThey are the restored gradient update value and the gradient output value respectively. The restored gradient update value is used to update the model weights in the model partitions configured for the current pipelined parallel group through an optimizer. The restored gradient output value is passed to the next parallel group in the reverse propagation direction for continued reverse propagation.

[0107] Calculate the gradient update value and the gradient output value according to the current update basis value, and determine whether the gradient update value is within a preset reasonable data range, which can effectively avoid the problems of gradient underflow and gradient explosion. If the gradient update value is reasonable, further perform gradient restoration calculation according to the gradient scaling value within the group to ensure the numerical stability of the gradient information. Using the restored gradient update value to update the model weights can more accurately adjust the model parameters and improve the training performance of the model. At the same time, passing the restored gradient output value to the next parallel group ensures the continuity and consistency of the gradient information in distributed training, thereby improving the utilization rate of distributed cluster resources and enhancing the efficiency, accuracy, and reliability of distributed training.

[0108] Further, based on the above embodiments, after determining whether the gradient update value is within a preset reasonable data range, it may further include:

[0109] If not, update the gradient scaling value within the group according to a preset gradient scaling value update rule, and trigger each pipelined parallel group located after the current pipelined parallel group in the reverse propagation direction to end the current round of model weight update operation.

[0110] Specifically, according to the current update basis value (i.e., the scaled gradient), calculate the gradient update value ▽W i ’ that matches the current pipelined parallel group, and calculate the gradient output value O i ’ that will be passed to the next parallel group. Check whether the gradient update value ▽W i ’ is within a reasonable range, such as checking whether the gradient is infinite, NaN (Not a Number), or exceeds a preset reasonable numerical range. If it is detected that a certain gradient update value overflows, the gradient scaling value of this parallel group needs to be adjusted, such as dividing the gradient scaling value by 2 to reduce the gradient scaling value of the current parallel group, or adopting other forms of update methods according to specific situations, such as multiplying the gradient scaling value by a factor less than 1 for reduction, to implement the update operation of the gradient scaling value within the current parallel group, and skip the gradient update link of the current parallel group to avoid training instability caused by numerical problems. Then, pass a failure message to the next parallel group to notify the subsequent parallel groups that there is a problem with the gradient calculation of the current parallel group, and the subsequent parallel groups can skip the gradient update operation of this round.

[0111] By adjusting the gradient scaling value when gradient overflow is detected and skipping the current gradient update step, numerical instability problems can be effectively avoided, ensuring numerical stability during the training process, avoiding gradient underflow and gradient explosion problems, and improving the robustness of distributed training. By dynamically adjusting the gradient scaling value, numerical problems can be flexibly addressed according to the actual gradient situation, thereby improving the adaptability of model weight updates, as well as the efficiency, accuracy, and reliability of distributed training.

[0112] Further, based on the above embodiments, after each of the model partitions is respectively configured in each pipelining parallel group, the following steps may further be included:

[0113] Input a plurality of quantized training samples into the configured quantization model for pre-training to obtain update jump information of the model weights in each of the model partitions;

[0114] According to the update jump information, determine the gradient scaling value update rules for each of the pipelining parallel groups configured with each of the model partitions.

[0115] Generally speaking, the update jump information can be specifically understood as: during the model training process, the change amplitude or change situation of the model weights at each update, reflecting the intensity of weight updates, that is, the change amount of the weights in one iteration, to evaluate the magnitude and stability of the gradient, thereby dynamically adjusting the gradient scaling value of the current parallel group.

[0116] Specifically, during the pre-training process, record the update situation of the model weights in each model partition, especially the jump amplitude of weight updates, and according to the weight update jump information of each model partition, determine the gradient scaling value update rules for each pipelining parallel group. When the gradient scaling value update rule is S i / a, for the pipelining parallel group with a relatively large jump amplitude of weight updates, a smaller initial scaling value S i can be set to reduce its jump amplitude and prevent gradient explosion. At the same time, set a larger update parameter a to significantly adjust its scaling value in subsequent training to prevent gradient explosion. For the pipelining parallel group with a relatively small jump amplitude of weight updates, set a larger initial scaling value S i to accelerate the propagation of the gradient and prevent gradient disappearance. At the same time, set a smaller update parameter a to finely adjust its scaling value in subsequent training to prevent gradient disappearance.

[0117] By obtaining the update jump information of model weights in each model partition through pre-training and dynamically adjusting the gradient scaling value update rule according to the information, the differences in data distribution among different parallel groups can be effectively addressed, avoiding the accuracy loss caused by the unified scaling value update rule, and significantly improving the convergence speed and final accuracy of the model in a distributed training environment. By setting different initial scaling values and update rules for different parallel groups, the complex requirements of large-scale model training can be better adapted, ensuring the stability, efficiency, and reliability of the training process.

[0118] For ease of understanding, the specific application scenarios applicable to each embodiment of the present disclosure will now be described. With the application of deep learning and the development of large model technologies, the number of model parameters continues to grow, and deep learning training consumes more and more time and resources. To improve computational efficiency and save computational resources, low-precision quantization techniques are generally used to accelerate model training. Low-precision quantization techniques specifically refer to using data types with smaller memory occupancy and lower precision to represent model parameters and computational inputs during training, thereby improving computational efficiency. However, the resulting accuracy loss problem may cause the model to fail to converge. To avoid this problem, error quantization techniques are usually further used to improve training accuracy.

[0119] Model training generally includes the following three main steps: forward calculation, reverse calculation of gradients, and parameter update using an optimizer. In the distributed training scenario involved in the present disclosure, model training is completed through the cooperation of multiple machines, and the key point is that multiple machines jointly responsible for the computational process of multiple consecutive layers in the neural network in a pipelined parallel manner. The machine combination that completes the same neural network computational layer is called a "pipelined parallel group". This distributed training method can distribute the computation of the model to each machine to complete in a relay manner, forming an efficient distributed computational process and significantly improving training efficiency.

[0120] Low-precision training uses a data type with lower precision (such as FP16) for calculation to complete the training process. Compared with using the default data type FP32 (32-bit floating-point number) for model training, FP16 has smaller memory occupancy and faster computational speed, but it will affect the model accuracy. To maintain the training accuracy of the model, additional operations need to be introduced for the reverse gradient calculation. After calculating the model error value of the reverse training data, multiply the error value by a fixed gradient scaling value, then perform the reverse gradient calculation. After the reverse gradient calculation is completed, divide by the scaling value to complete one round of reverse training of the neural network.

[0121] Figure 4It is a schematic diagram of a model training method using a unified scaling value provided by the related technology. Generally, before the training starts, the neural network model, the distributed training strategy, and the initial scaling value S are input. Then, according to the scale of the distributed training strategy, all the computing nodes participating in the calculation are evenly divided into n parallel groups pg 1 、pg 2 、…、pg n , and each parallel group contains the same number of computing nodes. At the same time, the neural network to be trained is evenly divided into n corresponding parallel groups layers 1 、layers 2 、…、layers n , and each parallel group contains the same number of neural network layers. During the forward training process, starting from the pg 1 parallel group, it receives the sample input of the model, calculates the output of layers 1 , and passes the data to the pg 2 parallel group. After the pg 2 parallel group receives the input, it calculates the output of layers 2 and passes it on. And so on until the pg n completes the neural network calculation of layers n to obtain the error value L of the neural network, thus completing the forward calculation process.

[0122] Before the backward calculation starts, first perform a scaling calculation on the error value to obtain a new error value L ’ = L × S, and use L ’ as the input for the backward calculation to perform the backward gradient calculation. The backward calculation process is similar to the forward calculation process, but it starts from the pg n to first calculate the gradient ▽W n n ’ of layers 1 1 of the last parallel group to be subjected to the backward calculation. After the pg 1 completes the calculation of the partial gradient ▽W ’ 2 、▽W ’ n 、...、▽W ’ ’ of layers1, first check whether there are overflow values in the gradients of ▽W ’ / S, and then use the scaled gradient as the parameter to update the input, completing the update of the model weights. Optionally, all of the above calculation processes can be performed in the FP16 data type.

[0123] In this technical solution where a unified scaling value is set for backpropagation training, for operators or calculation layers with different data distributions, this scaling value will become invalid, thus unable to effectively maintain the model accuracy and the convergence of the model. In addition, implementing corresponding error quantization schemes for each operator separately not only increases the difficulty of engineering implementation but also significantly improves the complexity of the entire system.

[0124] To solve the above problems, embodiments of the present disclosure propose a distributed training method for a neural network model. Figure 5 It is a schematic diagram of the overall execution process in the distributed training of a neural network model applicable to embodiments of the present disclosure. Figure 6 It is a schematic diagram of the execution process of a parallel group in the distributed training of a neural network model applicable to embodiments of the present disclosure. Specifically, this solution scales the backpropagation gradients at the granularity of training groups in the parallel training mode. Each parallel training group maintains its own independent scaling value. When receiving the gradients passed from the previous parallel group, it scales the received gradients. After all the backpropagation calculations in this parallel training group are completed, it restores the gradients to the original scaling and passes them to the next parallel group.

[0125] Before the training starts, each pipeline parallel group independently maintains the corresponding initial scaling value S i , the forward calculation process remains unchanged. When the backpropagation calculation starts, each parallel group receives the output O i+1 (non-first parallel group of the backpropagation calculation) or the unscaled error value L (first parallel group of the backpropagation calculation) from the previous parallel group, and then uses the unique scaling value S i of this parallel group to perform the scaling calculation O i+1 ’ = O i+1 × S i (non-first parallel group of the backpropagation calculation), or L ’ = L × S i (first parallel group of the backpropagation calculation), and then uses the scaled value to perform the backpropagation calculation. After the parallel group completes the backpropagation gradient calculation of the corresponding network layer layers i , check whether ▽W i ’ overflows. If it overflows, adjust the scaling value of this parallel group to S i / 2, then skip the gradient update step and input a failure message to the next communication group. If there is no numerical overflow, first scale the gradient ▽W i = ▽W i ’ / S i , and then use the scaled gradient as the parameter-updated input to complete the update of the model weights. And the output gradient O i = O i ’ / S i The scaling is passed to the previous communication group. The above forward calculation to weight update process is looped until the maximum number of training steps is reached. Among them, the scaling values that each parallel group needs to maintain are different, and the scaling value is a number greater than 1. Optionally, the scaling value can be 2 to the 20th power. The above calculation process can still be calculated using the FP16 calculation method.

[0126] Further, based on the above embodiments, before using the distributed training method of the neural network model proposed by the embodiments of the present disclosure, it is possible to detect whether the model scale of the model to be trained is greater than the processing scale of the distributed cluster, that is: according to the model scale of the model to be trained, predict the target storage requirement when training the model to be trained; obtain the standard storage size that the distributed cluster can provide; if the standard storage size does not meet the target storage requirement, it is determined that the model scale of the model to be trained is greater than the processing scale of the distributed cluster. When it is detected that the model scale of the model to be trained is greater than the processing scale of the distributed cluster, it means that the processing ability of the distributed cluster is insufficient to process this large-scale model, and it is necessary to use the distributed training method of the neural network model proposed by the embodiments of the present disclosure for distributed training. The method based on parallel groups makes the implementation of distributed training more modular, simplifies the complexity of engineering implementation, is particularly suitable for the distributed training scenarios of neural network models or large models, and improves the efficiency, accuracy and reliability of distributed training.

[0127] Optionally, before forward calculation, in addition to evenly dividing all participating nodes into n parallel groups according to the scale of the distributed training strategy (such as the total number of layers of each neural network layer and the preset basic partition number), and evenly dividing the neural network to be trained into n corresponding parallel groups. The following method can also be adopted: The model is pre-split into multiple alternative partitions according to the number of neural network layers; each alternative partition (parallel group) is pre-configured in each of the pipeline parallel groups, at least one quantized training sample is input into the pre-configured model for pre-testing, and the numerical distribution of each neural network layer in each pipeline parallel group is obtained. If it is determined that the numerical distribution of each neural network layer in at least one target pipeline parallel group does not meet the intra-group distribution consistency condition, then obtain the current numerical distribution of each current neural network layer in the current target pipeline parallel group, and according to each current numerical distribution, split each current neural network layer into multiple groups that meet the distribution consistency condition; where each group contains at least one current neural network layer; according to the position of each group in the model and the number value of the neural network layers included in each group, merge each group with adjacent alternative partitions.

[0128] That is, the general numerical distribution gap between adjacent neural network layers is similar. Due to the average distribution method, neural network layers that originally met the intra-group distribution consistency condition may be divided into different parallel groups, resulting in non-compliance with the intra-group distribution consistency condition. Therefore, the neural network layer that does not meet the intra-group distribution consistency condition can be used as the starting point for merging with adjacent neural network layers with similar forward or backward numerical distributions. If, after merging, the number of neural network layers in the parallel group is greater than the preset threshold, no merging is performed. Instead, the elements in the original parallel group that do not meet the intra-group distribution consistency condition are split so that the numerical distribution of the split parallel groups meets the intra-group distribution consistency condition. If there are multiple remaining groups after the merging process, then the multiple groups are respectively determined as different alternative partitions. Return to execute the operation of pre-configuring each alternative partition in each pipeline parallel group until the end iteration condition is met (meeting the convergence condition, such as the intra-group distribution consistency condition or the maximum number of iterations); determine each alternative partition obtained at the end of the iteration as each model partition matching the model. This ensures that the numerical distribution of the parallel groups meets the intra-group distribution consistency condition and the number of elements in the parallel groups is controlled within a reasonable range, guaranteeing the efficiency, accuracy, and reliability of distributed training.

[0129] Optionally, based on the above embodiments, when checking ▽W i ’ for overflow, in addition to adjusting the scaling value in a fixed manner, such as adjusting the scaling value of this parallel group to S i / 2. Additionally, respective update rules can be set for each pipelining parallel group. Specifically, after configuring each model partition in each pipelining parallel group, multiple training samples are input into the configured model for pre-training to obtain the update jump information of the model weights in each model partition. According to each update jump information, the gradient scaling value update rule S of each pipelining parallel group configured with each model partition is determined i / a. For example, for a pipelining parallel group with a relatively large update jump amplitude, a smaller initial scaling value S is set i , to reduce its jump amplitude and prevent gradient explosion, and a larger update parameter a is set to significantly adjust its scaling value. For a pipelining parallel group with a relatively small update jump amplitude, a larger initial scaling value S is set i , to accelerate the propagation of the gradient and prevent gradient vanishing, and a smaller update parameter a is set to finely adjust its scaling value. This can not only effectively handle the differences in data distribution between different parallel groups and avoid accuracy loss caused by uniformly adjusting the scaling value, but also significantly improve the convergence speed and final accuracy of the model in a distributed training environment, thus better adapting to the complex requirements of large-scale model training

[0130] The distributed training method of the neural network model proposed in the embodiments of the present disclosure is based on the training group in the parallel training mode as the granularity, and the scaling values that each parallel group needs to maintain are different values set according to business requirements. Each parallel group can effectively adapt to the unique numerical distribution within the parallel group to achieve better accuracy retention, and can be effectively combined with the parallel training mode. Since point-to-point network communication is required between each pipelining parallel group, a tensor multiplication is performed before the transmission of the reverse gradient, and the introduced additional calculation is very small, with low overhead, significantly improving the accuracy and stability of distributed training. The method based on the parallel group makes the implementation of distributed training more modular, simplifies the complexity of engineering implementation, is particularly suitable for the distributed training scenarios of neural network models or large models, and improves the efficiency, accuracy, and reliability of distributed training

[0131] As an implementation of the above-mentioned distributed training methods of each neural network model, the present disclosure also provides an optional embodiment of an execution device for implementing the above-mentioned distributed training methods of each neural network model

[0132] Figure 7 is a schematic diagram of a distributed training device for a neural network model provided according to an embodiment of the present disclosure, as shown in Figure 7 shown. The device includes: a quantization processing module 710, a training module 720, an error value generation module 730, and a weight update module 740, where

[0133] The quantization processing module 710 is configured to perform quantization processing on the to-be-trained model when it is detected that the model scale of the to-be-trained model is larger than the processing scale of the distributed cluster, so as to obtain a quantized model;

[0134] The training module 720 is configured to split the quantized model into multiple model partitions, and respectively configure each of the model partitions in each pipelined parallel group to perform distributed pipelined parallel training for the quantized model;

[0135] The error value generation module 730 is configured to generate a training error value through the joint cooperation among the pipelined parallel groups during the forward calculation of the quantized model according to the quantized training samples;

[0136] The weight update module 740 is configured to update the model weights in the configured model partitions through each of the pipelined parallel groups according to the gradient scaling values independently maintained within the group during the backward calculation of the quantized model according to the training error value.

[0137] Based on this, the technical solution of the embodiments of the present disclosure performs quantization processing on the model to obtain a quantized model when it is detected that the model scale of the to-be-trained model is larger than the processing scale of the distributed cluster; splits the quantized model into multiple model partitions and configures them in each pipelined parallel group; during the forward calculation, generates a training error value through the joint cooperation among the pipelined parallel groups; and during the backward calculation of the quantized model according to the training error value, updates the model weights through each of the pipelined parallel groups according to the gradient scaling values independently maintained within the group. Taking the parallel group as the granularity, each parallel group can effectively balance between the model accuracy and the computational overhead by maintaining its own gradient scaling value and using it to update the weights in its own model partition, improving the efficiency, accuracy, and reliability of distributed training.

[0138] Optionally, based on the above embodiments, the quantization processing module 710 is specifically configured to:

[0139] Predict the target storage requirement for training the to-be-trained model according to the model scale of the to-be-trained model;

[0140] Obtain the standard storage size that the distributed cluster can provide;

[0141] If the standard storage size does not meet the target storage requirement, it is determined that the model scale of the to-be-trained model is larger than the processing scale of the distributed cluster.

[0142] Optionally, based on the above embodiments, the training module 720 is specifically configured to:

[0143] Determine the number of neural network layers included in each model partition according to the total number of layers of each neural network layer included in the quantization model and a preset basic partition number.

[0144] Split the quantization model into multiple model partitions according to the number of neural network layers.

[0145] Optionally, based on the above embodiments, the training module 720 may include: a pre-splitting sub-module and a fine-tuning sub-module, where:

[0146] The pre-splitting sub-module is configured to pre-split the quantization model into multiple alternative partitions according to the number of neural network layers.

[0147] The fine-tuning sub-module is configured to start from the pre-split alternative partitions, perform at least one pre-test on the quantization model by using quantization training samples in combination, and perform at least one fine-tuning process on each of the alternative partitions according to the pre-test results to obtain the multiple model partitions that match the quantization model.

[0148] Optionally, based on the above embodiments, the fine-tuning sub-module may include: a pre-configuration unit, a pre-test unit, a fine-tuning unit, an operation return unit, and a model matching unit, where:

[0149] The pre-configuration unit is configured to pre-configure each of the alternative partitions in each of the pipelined parallel groups.

[0150] The pre-test unit is configured to input at least one quantization training sample into the quantization model after pre-configuration for pre-testing, and obtain the numerical distribution of each neural network layer in each of the pipelined parallel groups.

[0151] The fine-tuning unit is configured to, if it is determined that the numerical distribution of each neural network layer in at least one target pipelined parallel group does not satisfy the intra-group distribution consistency condition, perform fine-tuning on each of the alternative partitions according to the numerical distribution.

[0152] The operation return unit is configured to return the operation of pre-configuring each of the alternative partitions in each of the pipelined parallel groups until an end iteration condition is satisfied.

[0153] The model matching unit is configured to determine each of the alternative partitions obtained at the end of the iteration as each of the model partitions that match the quantization model.

[0154] Optionally, based on the above embodiments, the fine-tuning unit may include: a distribution acquisition sub-unit, a splitting sub-unit, a merging sub-unit, and a partition determination sub-unit, where:

[0155] A distribution acquisition subunit, configured to acquire the current numerical distributions of the current neural network layers in the current target pipelined parallel group;

[0156] A splitting subunit, configured to split each of the current neural network layers into multiple groups that meet the distribution consistency condition according to the current numerical distributions; wherein, each group contains at least one current neural network layer;

[0157] A merging subunit, configured to perform a merging process on each of the groups and adjacent alternative partitions according to the positions of the groups in the quantization model and the number values of the neural network layers included in each group;

[0158] A partition determination subunit, configured to, if the number of remaining groups after the merging process is multiple, respectively determine the multiple groups as different alternative partitions.

[0159] Optionally, based on the above embodiments, the weight update module 740 is specifically configured to:

[0160] If the current pipelined parallel group is the first parallel group in the backpropagation direction, perform a target scaling calculation on the training error value according to the gradient scaling value within the group to obtain the current update basis value;

[0161] If the current pipelined parallel group is not the first parallel group, perform a target scaling calculation on the gradient output value received from the previous pipelined parallel group in the backpropagation direction according to the gradient scaling value within the group to obtain the current update basis value;

[0162] Update the model weights in the configured model partition according to the current update basis value.

[0163] Optionally, based on the above embodiments, the weight update module 740 may include: a calculation sub-module, a judgment sub-module, a gradient reduction calculation sub-module, an update sub-module, and an output sub-module, wherein:

[0164] The calculation sub-module is configured to calculate a gradient update value and a gradient output value that match the current pipelined parallel group according to the current update basis value;

[0165] The judgment sub-module is configured to judge whether the gradient update value is within a preset reasonable data range;

[0166] The gradient reduction calculation sub-module is configured to, if so, perform a gradient reduction calculation on the gradient update value and the gradient output value according to the gradient scaling value within the group;

[0167] The update sub-module is configured to update the model weights in the configured model partition of the current pipelined parallel group according to the reduced gradient update value;

[0168] An output sub-module, configured to output the restored gradient output value to the next pipeline parallel group in the reverse propagation direction of the current pipeline parallel group.

[0169] Further, based on the above embodiments, the weight update module 740 may further include: an end sub-module, where:

[0170] The end sub-module is configured to, after determining whether the gradient update value is within a preset reasonable data range, if not, update the gradient scaling value within the group according to a preset gradient scaling value update rule, and trigger each pipeline parallel group after the current pipeline parallel group in the reverse propagation direction to end the current round of model weight update operation.

[0171] Further, based on the above embodiments, the training module 720 may further include: a jump information acquisition sub-module and a rule determination sub-module, where:

[0172] The jump information acquisition sub-module is configured to, after configuring each of the model partitions in each pipeline parallel group, input a plurality of quantization training samples into the configured quantization model for pre-training, and acquire the update jump information of the model weights in each of the model partitions;

[0173] The rule determination sub-module is configured to determine the gradient scaling value update rule of each of the pipeline parallel groups configured with each of the model partitions according to the respective update jump information.

[0174] The above product can execute the method provided in any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0175] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0176] According to an embodiment of the present disclosure, the present disclosure also provides a device, a readable storage medium, and a computer program product.

[0177] Figure 8FIG. 0 shows a schematic block diagram of an example device that may be used to implement embodiments of the present disclosure. The device refers to a chip including a GPU (Graphics Processing Unit) 810 or an NPU (Neural Processing Unit) 820, or both, and any suitable processor, controller, microcontroller, and other similar hardware components. These devices are specifically designed to efficiently execute computationally intensive tasks such as deep learning, graphics rendering, and data processing. They can be integrated into various computing systems, which are intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The computing system may operate in a distributed cluster including multiple pipelined parallel groups, each of which includes multiple devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0178] The device may execute the various methods and processes described above, such as the distributed training method of a neural network model, that is:

[0179] When it is detected that the model scale of the model to be trained is greater than the processing scale of the distributed cluster, perform quantization processing on the model to be trained to obtain a quantized model;

[0180] Split the quantized model into multiple model partitions, and respectively configure each model partition in each pipelined parallel group to perform distributed pipelined parallel training for the quantized model;

[0181] During the forward calculation of the quantized model according to the quantized training samples, generate a training error value through the cooperation among the pipelined parallel groups;

[0182] During the backward calculation of the quantized model according to the training error value, update the model weights in the configured model partitions by each pipelined parallel group according to the gradient scaling value independently maintained within the group.

[0183] For example, in some embodiments, the distributed training method of a neural network model may be implemented as a computer software program, which is tangibly included in a machine-readable medium. In some embodiments, part or all of the computer program may be loaded and / or installed on the device. When the computer program is executed by the device, one or more steps of the distributed training method of the neural network model described above may be executed.

[0184] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0185] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0186] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0187] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0188] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0189] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system, or a server combined with a blockchain.

[0190] Artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.

[0191] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, storage devices, etc., and the resources can be deployed and managed in a demand-driven and self-service manner. Through cloud computing technology, it is possible to provide efficient and powerful data processing capabilities for the application of technologies such as artificial intelligence and blockchain and model training.

[0192] Figure 9 FIG. 4 is a schematic diagram of a distributed cluster provided according to an embodiment of the present disclosure. As Figure 9 shown, the distributed cluster includes multiple devices such as device 1910, device 2920, device 3930, …, device N 940 described in the above embodiments. Each device is used to execute the various methods and processes described above, such as the distributed training method of a neural network model, that is:

[0193] When it is detected that the model scale of the model to be trained is greater than the processing scale of the distributed cluster, quantization processing is performed on the model to be trained to obtain a quantized model;

[0194] The quantized model is split into multiple model partitions, and each of the model partitions is respectively configured in each pipelined parallel group to perform distributed pipelined parallel training for the quantized model;

[0195] During the forward calculation of the quantized model based on the quantization training samples, a training error value is generated through the joint cooperation between the pipelined parallel groups;

[0196] During the backward calculation of the quantized model based on the training error value, the model weights in the configured model partitions are updated by each of the pipelined parallel groups according to the gradient scaling values independently maintained within the group.

[0197] In the distributed cluster according to the embodiment of the present disclosure, each device in the cluster works together to jointly execute the distributed training task of the neural network model. When the device in the cluster detects that the model scale of the model to be trained is greater than the processing scale of the distributed cluster, quantization processing is performed on the model to obtain a quantized model; the quantized model is split into multiple model partitions and configured in each pipelined parallel group; during the forward calculation, a training error value is generated through the joint cooperation between the pipelined parallel groups; during the backward calculation of the quantized model based on the training error value, the model weights are updated by each of the pipelined parallel groups according to the gradient scaling values independently maintained within the group. Taking the parallel group as the granularity, each parallel group can effectively balance between model accuracy and computational overhead by maintaining its own gradient scaling value and using it to update the weights in its own model partition, thereby improving the efficiency, accuracy, and reliability of distributed training.

[0198] Optionally, based on the above embodiments, the devices in the distributed cluster are computing chips.

[0199] Specifically, by configuring computing chips in the distributed cluster, when performing distributed training of a neural network model, the high parallel processing ability and efficient memory management of the chips can be fully utilized, reducing the communication overhead of distributed training. By splitting the quantization model into multiple model partitions and configuring them in each pipelined parallel group, the parallel computing ability of the computing chips can be fully utilized, significantly improving the speed of model training. The low-precision quantization technology further reduces the overhead of data transmission and computing, enabling the computing chips to more efficiently process the training tasks of large-scale models. Each parallel group independently maintains the gradient scaling value, avoiding the overhead of global synchronization, further optimizing the utilization efficiency of computing resources, and being able to better adapt to the numerical distribution of different partitions, thereby improving the overall accuracy of the model, effectively reducing the accuracy loss caused by low-precision quantization, and ensuring the convergence performance of the model during distributed training. The pipelined parallel method reduces the burden on a single computing chip by processing different parts of the model in stages, avoiding training instability caused by overloading of a single device. The independent maintenance and update mechanism of the gradient scaling value further improves the stability of the training process. The method based on the parallel group as the granularity makes the implementation of distributed training more modular, simplifies the complexity of engineering implementation, can effectively balance the model accuracy and computing overhead, and improves the efficiency, accuracy, and reliability of distributed training.

[0200] Optionally, based on the above embodiments, the computing chip may include a graphics processing unit or an artificial intelligence processor.

[0201] Specifically, by using a computing chip of a Graphics Processing Unit (GPU) or an Artificial Intelligence Processor (NPU), the distributed cluster can, when performing complex computing tasks (such as distributed training of deep learning or large models), split the model into multiple partitions and allocate them to different parallel groups. The parallel processing capabilities of the GPU and NPU can efficiently process low-precision data types, reduce memory occupancy and computing overhead, improve computing efficiency, reduce computing time, accelerate model training, and at the same time reduce precision loss through error quantization technology. By splitting the model into multiple partitions and allocating them to different parallel groups, the parallel capabilities of the computing chip can be fully utilized to accelerate the training of large-scale models. In addition, the low-power design of the NPU makes it suitable for use in mobile devices and edge computing scenarios, further expanding the application scope of the computing chip. This not only improves the overall performance of the system, but also optimizes power consumption and resource utilization, enabling the computing system to be more efficient at the granularity of parallel groups. Each parallel group maintains its own gradient scaling value and uses it to update the weights in its respective model partition, achieving a balance between model accuracy and computing overhead in the distributed training of the neural network model, and improving the efficiency, accuracy, and reliability of distributed training.

[0202] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution provided in this disclosure can be achieved. No limitation is made herein.

[0203] The above specific embodiments do not limit the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A distributed training method for a neural network model, executed by a distributed cluster, wherein the distributed cluster includes a plurality of pipeline parallel groups, each pipeline parallel group includes a plurality of devices, and the method comprises: When it is detected that the model scale of the to-be-trained model is larger than the processing scale of the distributed cluster, quantizing the to-be-trained model to obtain a quantized model; Splitting the quantized model into a plurality of model partitions, and configuring each of the model partitions in each pipeline parallel group, so as to perform distributed pipeline parallel training for the quantized model; In the process of forward calculation of the quantization model according to the quantization training samples, a training error value is generated through the cooperation between the pipeline parallel groups; In the process of reversely calculating the quantization model according to the training error value, the model weights in the configured model partitions are updated by each of the pipeline parallel groups according to the gradient scaling values ​​independently maintained within the group.

2. The method according to claim 1, wherein: The detecting that the model scale of the to-be-trained model is larger than the processing scale of the distributed cluster includes: Predicting target storage requirements for training the model to be trained based on the model size of the model to be trained; Obtaining a standard storage size that can be provided by the distributed cluster; If the standard storage size does not meet the target storage requirement, it is determined that the model scale of the to-be-trained model is larger than the processing scale of the distributed cluster.

3. The method according to claim 1, wherein: The step of splitting the quantization model into a plurality of model partitions comprises: Determine the number of neural network layers included in each model partition according to the total number of layers of each neural network layer included in the quantization model and the preset number of basic partitions; The quantization model is split into multiple model partitions according to the number of neural network layers.

4. The method according to claim 3, wherein: The step of splitting the quantization model into a plurality of model partitions according to the number of neural network layers specifically includes: Pre-splitting the quantization model into a plurality of candidate partitions according to the number of neural network layers; Taking the pre-split candidate partitions as the starting point, the quantization model is pre-tested at least once in combination with quantization training samples, and each of the candidate partitions is fine-tuned at least once based on the pre-test results to obtain the multiple model partitions that match the quantization model.

5. The method according to claim 4, wherein: The method of taking the pre-split candidate partitions as a starting point, pre-testing the quantization model at least once in combination with quantization training samples, and fine-tuning each of the candidate partitions at least once according to the pre-test results to obtain the multiple model partitions matching the quantization model includes: Pre-configuring each of the candidate partitions in each of the pipeline parallel groups; Input at least one quantized training sample into the pre-configured quantized model for pre-testing, and obtain the numerical distribution of each neural network layer in each pipeline parallel group; If it is determined that the value distribution of each neural network layer in at least one target pipeline parallel group does not meet the intra-group distribution consistency condition, then fine-tuning each of the candidate partitions according to the value distribution; Returning to execute the operation of pre-configuring each candidate partition in each pipeline parallel group, until the iteration termination condition is met; The candidate partitions obtained when the iteration ends are determined as the model partitions that match the quantization model.

6. The method according to claim 5, wherein: The fine-tuning of each candidate partition according to the value distribution specifically includes: Get the current value distribution of each current neural network layer in the parallel group with the current target pipeline; According to each of the current numerical distributions, each of the current neural network layers is split into a plurality of groups that meet a distribution consistency condition; wherein each group contains at least one current neural network layer; According to the position of each group in the quantization model and the number of neural network layers contained in each group, each group is merged with an adjacent candidate partition; If there are multiple groups remaining after the merging process is completed, the multiple groups are respectively determined as different candidate partitions.

7. The method according to any one of claims 1 to 6, wherein: In the process of reversely calculating the quantization model according to the training error value, updating the model weights in the configured model partitions according to the gradient scaling values ​​independently maintained in the groups by each of the pipeline parallel groups, including: If the current pipeline parallel group is the first parallel group in the reverse propagation direction, the training error value is subjected to target scaling calculation according to the gradient scaling value in the group to obtain the current update basis value; If the current pipeline parallel group is not the first parallel group, then according to the gradient scaling value in the group, the target scaling calculation is performed on the gradient output value received from the previous pipeline parallel group in the reverse propagation direction to obtain the current update basis value; The model weights in the configured model partitions are updated according to the current update basis value.

8. The method according to claim 7, wherein: The updating of the model weights in the configured model partitions according to the current update basis value includes: Calculate the gradient update value and the gradient output value that match the current pipeline parallel group according to the current update basis value; Determine whether the gradient update value is within a preset reasonable data range; If yes, performing gradient restoration calculation on the gradient update value and the gradient output value according to the gradient scaling value in the group; Update the model weights in the configured model partitions of the current pipeline parallel group according to the restored gradient update values; The restored gradient output value is output to the subsequent pipeline parallel group of the current pipeline parallel group in the reverse propagation direction.

9. The method according to claim 8, after determining whether the gradient update value is within a preset reasonable data range, further comprises: If not, the gradient scaling value in the group is updated according to the preset gradient scaling value update rule, and each pipeline parallel group located after the current pipeline parallel group in the reverse propagation direction is triggered to end the update operation of the model weight of this round.

10. The method according to claim 9, after configuring each of the model partitions in each pipeline parallel group, further comprises: Inputting a plurality of quantization training samples into the configured quantization model for pre-training, and obtaining update jump information of the model weights in each of the model partitions; According to each of the update jump information, the gradient scaling value update rule of each of the pipeline parallel groups configured with each of the model partitions is determined.

11. A distributed training device for a neural network model, configured in a distributed cluster, wherein the distributed cluster includes a plurality of pipeline parallel groups, each pipeline parallel group includes a plurality of devices, and the device includes: A quantization processing module, configured to quantize the model to be trained to obtain a quantized model when it is detected that the model scale of the model to be trained is larger than the processing scale of the distributed cluster; A training module, used for splitting the quantization model into a plurality of model partitions, and respectively configuring each of the model partitions in each pipeline parallel group, so as to perform distributed pipeline parallel training for the quantization model; An error value generating module, used for generating a training error value by cooperating between the pipeline parallel groups during the process of forward calculation of the quantization model according to the quantization training samples; The weight updating module is used to update the model weights in the configured model partitions according to the gradient scaling values ​​independently maintained within each pipeline parallel group during the process of reversely calculating the quantization model according to the training error value.

12. A device for executing the method according to any one of claims 1 to 10.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause a device to execute a method according to any one of claims 1-10.

14. A computer program product, comprising a computer program, wherein when the computer program is executed by a device, the steps of the method according to any one of claims 1 to 10 are implemented.

15. A distributed cluster comprising a plurality of devices, each device being configured to execute the method according to any one of claims 1 to 10.

16. The distributed cluster according to claim 15, wherein: The device is a computing chip.

17. The distributed cluster according to claim 16, wherein: The computing chip includes a graphics processing unit or an artificial intelligence processor.

Citation Information

Cited By

  • Neural network distributed training method, system and device based on mesh assembly line

    CN121212199A

  • Mesh-pipelined neural network distributed training method, system and device

    CN121212199B