Parallel training method and device for neural network model

The method optimizes neural network training by dividing the model and data into pipeline stages with tailored storage policies, addressing resource management and training time challenges in parallel training.

JP2025097908APending Publication Date: 2025-07-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024197870
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-11-13
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing neural network training methods face challenges in efficiently utilizing resources and reducing training time, particularly in managing memory usage and recomputation during parallel training processes.

Method used

A method and apparatus for parallel training of neural networks that divide the model and data set into pipeline stages, employing individual storage policies for activations generated by forward propagation to optimize memory usage and reduce recomputation.

Benefits of technology

This approach enhances training efficiency by minimizing memory usage and reducing training time while maintaining effective parallel training strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025097908000001_ABST
    Figure 2025097908000001_ABST
Patent Text Reader

Abstract

To provide a parallel training method and device for a neural network model.SOLUTION: This method includes steps for: dividing, based on a pipeline stage for a parallel training for a neural network model, the neural network model and a training data set; determining a partial model of the neural network model and a partial data set of the training data set; selecting individually a preservation policy for an activation generated by a forward propagation of the partial model to be utilized in order to calculate a gradient of a back propagation of the partial model with respect to a corresponding data set in the partial data set processed in the pipeline stage; and generating a strategy for the parallel training for the neural network model based on the preservation policy.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The following embodiments relate to a method and an apparatus for parallel training of a neural network model.

Background Art

[0002] The technical automation of the recognition process is realized, for example, via a neural network model embodied as a processor with a special calculation structure, which provides an intuitive calculation mapping between an input pattern and an output pattern after considerable training. The trained ability to generate such a mapping can be said to be the learning ability of the neural network. Moreover, through specialized training, such a specially trained neural network has, for example, a generalization ability to generate relatively accurate outputs for input patterns that have not been trained. When attempting to process operations related to the training and inference of such a neural network model, schemes such as model parallelization and / or data parallelization can be used as a solution that can converge to the result more quickly.

Summary of the Invention

Problems to be Solved by the Invention

[0003] An object of the present invention is to provide a method and an apparatus for parallel training of a neural network model.

Means for Solving the Problems

[0004] According to one embodiment, a parallel training method divides a neural network model and a training data set based on pipeline stages for parallel training of the neural network model to determine a partial model of the neural network model and a partial data set of the training data set, and determines, for use in calculating the gradient of backpropagation of the partial model, a storage policy for activations generated by forward propagation of the partial model for the corresponding data set of the partial data set processed in the pipeline stage. The method includes individually selecting the storage policy and generating a strategy for parallel training of the neural network model based on the storage policy.

[0005] An electronic device according to one embodiment includes a processor and a memory storing instruction codes. When the instruction codes are executed by the processor, the electronic device divides a neural network model and a training data set based on pipeline stages for parallel training of the neural network model to determine a partial model of the neural network model and a partial data set of the training data set, individually selects, for use in calculating the gradient of backpropagation of the partial model, a storage policy for activations generated by forward propagation of the partial model for the corresponding data set of the partial data set processed in the pipeline stage, and can generate a strategy for parallel training of the neural network model based on the storage policy.

[0006] A method for training a neural network model according to an embodiment includes: calculating a first value in a first pipeline stage of the neural network model based on a first partial dataset; processing the first value according to a first storage policy based on the first pipeline stage and the first partial dataset; calculating a second value in the first pipeline stage of the neural network model based on a second partial dataset; processing the second value according to a second storage policy based on the first pipeline stage and the second partial dataset; calculating a third value in a second pipeline stage of the neural network model based on the first partial dataset; processing the third value according to a third storage policy based on the second pipeline stage and the first partial dataset; calculating a fourth value in the second pipeline stage of the neural network model based on the second partial dataset; and processing the fourth value according to a fourth storage policy based on the second pipeline stage and the second partial dataset.

Advantages of the Invention

[0007] According to the present invention, a method and an apparatus for parallel training of a neural network model can be provided.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying Out the Invention

[0009] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and can be changed into various forms. Therefore, the embodiments are not limited to specific disclosed forms, and the scope of this specification includes changes, equivalents or alternatives included in the technical idea.

[0010] Terms such as first or second may be used to describe a plurality of components, but such terms must be interpreted only for the purpose of distinguishing one component from other components. For example, the first component can be named the second component, and similarly, the second component can also be named the first component.

[0011] When it is mentioned that any component is "connected" or "coupled" to another component, it should be understood that it is directly connected or coupled to the other component, but other components may exist in the middle.

[0012] Singular expressions include plural expressions unless the context clearly indicates otherwise. As used herein, terms such as "comprising" or "having" are intended to indicate the presence of the features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should not be construed as precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0013] As used herein, phrases such as "at least one of A or B" and "at least one of A, B, or C" can each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof.

[0014] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which this embodiment belongs. Commonly used pre-defined terms should be construed to have a meaning consistent with their meaning in the context of the relevant art, and should not be construed as having an ideal or overly formal meaning unless clearly defined herein.

[0015] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. When describing with reference to the drawings, the same components will be given the same reference numerals regardless of the reference signs, and redundant descriptions thereof will be omitted.

[0016] FIG. 1 is a diagram exemplarily showing a pipeline parallelization operation without recomputation according to an embodiment. Referring to FIG. 1, parallel training of a neural network model is performed using pipeline stages 110 to 140. The pipeline stages 110 to 140 can be performed by different processors. For example, the processor may include one or more of a GPU (graphic processing unit), an NPU (neural processing unit), and an accelerator. For example, the first to fourth pipeline stages 110 to 140 may be performed by the first to fourth GPUs. Although four pipeline stages 110 to 140 are illustrated in FIG. 1, the present invention is not limited thereto.

[0017] The neural network model and the training dataset can be divided based on the pipeline stages 110 to 140. For example, the neural network model may be divided into partial models, and the training dataset may be divided into partial datasets. For example, the layers of the neural network model may be divided into partial models, or the weights of the neural network model may be divided into partial models. The partial dataset corresponds to micro-batches. Hereinafter, typically, an example in which the partial model includes partial layers of the neural network model will be described, but the present invention is not limited thereto.

[0018] The partial model can be trained in pipeline stages 110 to 140. For example, the neural network model may be divided into a first partial model to a fourth partial model, and the first partial model to the fourth partial model may be trained in pipeline stages 110 to 140. For example, the first partial model to the fourth partial model may be trained in pipeline stages 110 to 140 by the first GPU to the fourth GPU.

[0019] The partial dataset can be trained in pipeline stages 110 to 140. For example, the training dataset may be divided into a first partial dataset to an eighth partial dataset. The partial datasets may be trained sequentially and in parallel using a pipeline method. For example, the first corresponding dataset of the first partial dataset may be trained in the first pipeline stage 110. Next, based on the training operation result of the first corresponding dataset of the first partial dataset in the first pipeline stage 110, the second corresponding dataset of the first partial dataset may be trained in the second pipeline stage 120.

[0020] Based on the training result, the first partial dataset can be represented in other states in pipeline stages 110 to 140. The partial dataset of each state can be called like the corresponding dataset. When the i-th partial dataset is trained in the j-th pipeline stage 110, the i-th partial dataset can be called the j-th corresponding dataset of the i-th partial dataset. i is a natural number from 1 to I, and j is a natural number from 1 to J. In the example shown in FIG. 1, I may be 8 and J may be 4. 1 to 8 in FIG. 1 indicate the i values of the partial datasets.

[0021] The training of a neural network model includes forward propagation and backward propagation. The training operation result means the activation by forward propagation. When the forward propagation using each partial dataset is completed, the backward propagation using each partial dataset is performed. For example, referring to the training process of the partial dataset identified as the value of 1 in FIG. 1, the forward propagation based on the first partial dataset is sequentially performed in the first pipeline stage 110 to the fourth pipeline stage 140. When the forward propagation based on the first partial dataset is completed, the backward propagation based on the first partial dataset is sequentially performed in the fourth pipeline stage 140 to the first pipeline stage 110.

[0022] The neural network model may be trained using gradients. Activation is generated by forward propagation, and gradients are calculated using the activation in backward propagation. The activation generated during forward propagation is kept in memory for use when calculating the gradients. All or part of the activation is stored in memory. Since the calculation of the activation is not separately required when calculating the gradients, the training time can be shortened. Since the memory space is used by the activation until the gradients are calculated, there is a risk of memory shortage if the size of the activation generated during forward propagation is large.

[0023] Instead of a policy of storing all or part of the activations in memory to calculate gradients, a policy of recomputing the activations when calculating gradients may be used. In this case, if an activation is generated during forward propagation, the memory space occupied by the activation is first released. The released memory space may be used by other data. If an activation is required for calculating gradients during backpropagation, the activation may be newly generated through recomputation. Since the memory space occupied by the activation is minimized, a memory shortage can be prevented. Recomputing the activation when calculating gradients may increase the training time.

[0024] According to an embodiment, a keeping policy of activations generated by forward propagation of a sub-model for use in calculating gradients of backpropagation of the sub-model of a neural network model can be individually selected for a corresponding dataset of the partial dataset processed in pipeline stages 110 to 140. For example, a first keeping policy is used to train the corresponding dataset of the first partial dataset of the partial dataset in the first pipeline stage 110, and a second keeping policy different from the first keeping policy may be used to train the corresponding dataset of the second partial dataset of the partial dataset in the first pipeline stage 110. According to an embodiment, the keeping policy can be selected based on the memory usage status. The keeping policy may be selected so as to maximize memory usage within the available memory capacity.

[0025] Once the selection of the storage policy is completed, a strategy for parallel training of the neural network model is generated based on the selected storage policy. For example, the corresponding strategy may include information about pipeline stages 110-140, sub-models of the neural network model, sub-datasets of the training dataset, the storage policy, or a combination thereof. The neural network model can be trained based on the corresponding strategy.

[0026] FIG. 2 is a diagram showing the training process of a neural network model using gradients according to an embodiment. Referring to FIG. 2, the neural network model 210 includes an (n-2)th layer 211, an (n-1)th layer 212, and an nth layer 213. Layers 211, 212, and 213 each include weights, and the weights can be updated during the training process. The weights may also be referred to as parameters. n indicates the number of layers of the neural network model 210. FIG. 2 shows an example where n = 3, but is not limited thereto.

[0027] For the training of the neural network model 210, forward propagation operations 221, 222, 223, operations 231, 232, 233 for calculating the activation gradient, operations 241, 242, 243 for calculating the layer gradient, and model update operations 251, 252, 253 are performed. The operations 231, 232, 233 for calculating the activation gradient and the operations 241, 242, 243 for calculating the layer gradient correspond to the backpropagation operation.

[0028] The forward propagation operation 221 executes the (n-2)th layer 211 based on the input data I and generates the hidden activation H1. The forward propagation operation 222 executes the (n-1)th layer 212 based on the hidden activation H1 and generates the hidden activation H2. The forward propagation operation 223 executes the nth layer 213 based on the hidden activation H2 and generates the output activation O.

[0029] Activation gradient data AG1 is determined according to the comparison result between the output activation O and the ground truth GT. The activation gradient data AG1 corresponds to a loss. The arithmetic operation 231 generates activation gradient data AG2 using the activation gradient data AG1 and the n-th layer 213, the arithmetic operation 232 generates activation gradient data AG3 using the activation gradient data AG2 and the (n - 1)-th layer 212, and the arithmetic operation 233 generates activation gradient data AG4 using the activation gradient data AG3 and the (n - 2)-th layer 211.

[0030] The arithmetic operation 241 generates layer gradient data LG1 using the activation gradient data AG1 and the hidden activation H2, the arithmetic operation 242 generates layer gradient data LG2 using the activation gradient data AG2 and the hidden activation H1, and the arithmetic operation 243 generates layer gradient data LG3 using the activation gradient data AG3 and the input data I. The gradient is indicated by ΔW. W indicates a weight. The well-known term "gradient" refers to the layer gradient. The layer gradient is simply referred to as the gradient.

[0031] Thus, in the backpropagation operation, the activations H1, H2, O generated by the forward propagation operations 221, 222, 223 are required for calculating the layer gradient. According to the embodiment, all or part of the activations H1, H2, O can be stored in the memory according to the storage policy.

[0032] The model update operation 251 updates the weights of the n-th layer 213 based on the layer gradient data LG1. The model update operation 252 updates the weights of the (n - 1)-th layer based on the layer gradient data LG2. The model update operation 253 updates the weights of the (n - 2)-th layer 211 based on the layer gradient data LG3. The model update operations 251, 252, and 253 are represented as W-(γ*ΔW). γ represents the learning rate. The trained n-th layer 261, the trained (n - 1)-th layer 262, and the trained (n - 2)-th layer 263 are generated by the model update operations 251, 252, and 253.

[0033] FIG. 3 is a graph exemplarily showing the memory usage of each pipeline stage in the case of no recomputation according to an embodiment. Referring to FIG. 3, it can be seen that the memory usage gradually decreases as it progresses from the first pipeline stage (e.g., the first pipeline stage 110 shown in FIG. 1) to the final pipeline stage (e.g., the fourth pipeline stage 140 shown in FIG. 1). According to the pipeline method, the closer the pipeline stage is to the first pipeline stage, the greater the interval between the time when the forward propagation is performed and the time when the backward propagation is performed, and the more activations that need to be stored from the time when the backward propagation is completed to the time when the backward propagation is performed. According to the embodiment, according to such memory usage characteristics, for each corresponding dataset of each pipeline stage, the storage policy of the activations can be individually selected.

[0034] FIG. 4 is a diagram exemplarily showing an activation storage policy according to an embodiment. Referring to FIG. 4, forward propagation of corresponding data sets 1 to 4 of the first to fourth partial data sets is performed in the first pipeline stage 410. Through the forward propagation, activations of the corresponding data sets 1 to 4 are generated. The activations are stored in the memory. Depending on the storage policies 411 to 413 of the corresponding activations, the corresponding activations may be stored in the memory until backpropagation, or the memory space in which the corresponding activations are stored may be immediately released without storing the corresponding activations.

[0035] The storage policies 411 to 413 can be selected based on the memory usage state. The storage policies 411 to 413 may be selected so as to maximize the memory usage within the available memory capacity. The storage policies 411 to 413 include a full recomputation policy that generates an activation by performing recomputation during backpropagation without storing the activation, a partial recomputation policy that stores a part of the activation and generates the remaining part of the activation by performing partial recomputation during backpropagation, and a non-recomputation policy that stores the activation and does not perform recomputation and partial recomputation during backpropagation, and may include one or more of them.

[0036] FIG. 5 is a diagram exemplarily showing the structure of a neural network model according to an embodiment. Referring to FIG. 5, the neural network model 500 includes layers 510 to 530, and the layers 510 to 530 include various operations. The corresponding operations may be grouped according to various criteria. For example, the k-th layer 520 may include operation groups 521 to 523, and the n-th layer 530 may include operation groups 531 and 532. The amounts of operations of the operation groups 521 to 523, 531, and 532 and the capacities of the operation results of the operation groups 521 to 523, 531, and 532 may vary. Based on the characteristics of such operation groups 521 to 523, 531, and 532, a partial recomputation policy is defined.

[0037] For example, the fourth operation group 531 may have a smaller amount of operations than the fifth operation group 532, but the operation result of the fourth operation group 531 may occupy a larger capacity than the operation result of the fifth operation group 532. In this case, it is efficient to use the fourth operation group 531 as the partial recomputation target of the partial recomputation policy. This is because the memory space efficiency can be increased and the time delay due to recomputation can be minimized. The k-th layer 520 and the n-th layer 530 have different internal structures together with the operation groups 521 to 523 and the operation groups 531 and 532. Different operation groups in the k-th layer 520 and the n-th layer 530 may be used as partial recomputation targets.

[0038] The neural network model 500 may include layers having the same internal structure. For example, the layer may be a transformer layer. The transformer layer may include a self-attention operation group and an MLP (multi-layer perceptron) operation group. The self-attention operation group has a smaller amount of operations compared to the MLP operation group, and the operation result of the self-attention group has a larger capacity compared to the operation result of the MLP operation group. In the transformer layer, the self-attention operation group may be used as the subject of partial recomputation of the partial recomputation policy.

[0039] For example, the layer of any pipeline stage in the pipeline stage may have the same structure as the k-th layer 520, and the layer of any other pipeline stage may have the same structure as the n-th layer 530. For example, the layer of the first pipeline stage may have the same structure as the n-th layer 530. In this case, any operation group of the layer of the first pipeline stage can be selected as the subject of partial recomputation. For example, taking the case where the fourth operation group 531 is selected as the subject of partial recomputation. In this case, after the forward propagation of the layer of the first pipeline stage is performed, the memory space storing the activation of the fourth operation group 531 is released. The activation of the fifth operation group 532 may be stored for gradient calculation in backpropagation. When the backpropagation of the layer of the first pipeline stage is performed, partial recomputation of the fourth operation group 531 is performed to calculate the gradient. The activation of the fifth operation group 532 is loaded from the memory without separate recomputation. For example, the n-th layer 530 may be a transformer layer, the fourth operation group 531 may be a self-attention operation group, and the fifth operation group 532 may be an MLP operation group.

[0040] FIG. 6 is a diagram exemplarily showing a pipeline parallelization operation including recomputation according to an embodiment. Referring to FIG. 6, in the first training process 610, a non-recomputation policy is selected for the first corresponding data set of the first partial data set in the first pipeline stage 611, the second corresponding data set of the first partial data set in the second pipeline stage 612, the third corresponding data set of the first partial data set in the third pipeline stage 613, the third corresponding data set of the second partial data set in the third pipeline stage 613, and the fourth corresponding data set of the first partial data set in the fourth pipeline stage 614, and a full recomputation policy is selected for the first corresponding data set of the second partial data set in the first pipeline stage 611, the first corresponding data set of the third partial data set in the first pipeline stage 611, the first corresponding data set of the fourth partial data set in the first pipeline stage 611, the second corresponding data set of the second partial data set in the second pipeline stage 612, and the second corresponding data set of the third partial data set in the second pipeline stage 612.

[0041] In the forward operation of the first training process 610, the activations of the corresponding data sets to which the non-recomputation policy is applied are stored in the memory, and the activations of the corresponding data sets to which the full recomputation policy is applied may also be stored in the memory. In the backward operation of the first training process 610, the activations of the corresponding data sets to which the non-recomputation policy is applied are loaded from the memory for gradient calculation, and the activations of the corresponding data sets to which the full recomputation policy is applied can be newly generated for gradient calculation.

[0042] The second training process 620 is an example in which a partial recomputation policy is selected for the second corresponding data set of the second partial data set in the second pipeline stage 612 and the second corresponding data set of the third partial data set in the second pipeline stage 612, which is different from the first training process 610.

[0043] In the forward operation of the second training process 620, the activations of the corresponding data sets to which the non-recomputation policy is applied are stored in memory, the activations of the corresponding data sets to which the full recomputation policy is applied may not be stored in memory, and the activations of the corresponding data sets to which the partial recomputation policy is applied may be partially stored in memory. In the backward operation of the second training process 620, the activations of the corresponding data sets to which the non-recomputation policy is applied are loaded from memory for gradient calculation, the activations of the corresponding data sets to which the full recomputation policy is applied are newly generated for gradient calculation, and the activations of the corresponding data sets to which the partial recomputation policy is applied may have a part loaded from memory and the remaining part newly generated for gradient calculation.

[0044] Unlike the first training process 610, the second training process 620 may optionally apply the partial recomputation policy. A part of the activations of the corresponding data sets to which the partial recomputation policy is applied may be loaded from memory without separate recomputation. Therefore, the memory usage level of the second training process 620 is higher than that of the second training process 620, and the training time of the second training process 620 can be shortened by the time that the recomputation is reduced by using the partial recomputation policy compared to the first training process 610. The illustration in FIG. 6 is for the sake of convenience of explanation and relates to a limited number of pipeline stages 611 to 614, 621 to 624 and a limited number of partial data sets, but a larger number of pipeline stages and a larger number of partial data sets are used than in reality, and in this case, the training time can be further significantly shortened.

[0045] FIG. 7 is a flowchart exemplarily showing a process of selecting a storage policy using priorities according to an embodiment. Referring to FIG. 7, in step 710, the electronic device calculates the memory usage of activation for each storage policy regarding the corresponding data set of the partial data set. The electronic device can calculate the memory usage of activation for each storage policy of the corresponding data set by simulating the training process of each partial model.

[0046] For example, the storage policy may include one or more of a full recomputation policy, a partial recomputation policy, and a non-recomputation policy. When one or more of the full recomputation policy, the partial recomputation policy, and the non-recomputation policy are applied to each corresponding data set, the electronic device can calculate the capacity occupied in the memory until backpropagation is performed after the activation of each corresponding data set is generated. When full recomputation is used, the activation does not occupy a separate memory space until backpropagation. When non-recomputation is used, the entire activation acquires a memory space until backpropagation. When partial recomputation is used, a part of the activation (for example, the part generated by the operation group that does not perform recomputation) occupies the memory space until backpropagation.

[0047] In step 720, the electronic device calculates the available memory capacity for each corresponding data set for each pipeline stage. The memory usage level increases when the training process using the corresponding data set is performed. The available memory capacity when storing the activation in the memory according to each storage policy can be calculated.

[0048] In step 730, the electronic device selects an activation storage policy for each corresponding data set based on the priority. The non-recomputation policy has a higher priority than the partial recomputation policy, and the partial recomputation policy has a higher priority than the full recomputation policy. For example, according to the memory usage status, if all of the full recomputation policy, partial recomputation policy, and non-recomputation policy are available for any corresponding data set, the non-recomputation policy can be selected.

[0049] If there is competition by priority among the training operations of the corresponding data sets, the priority can be applied earlier to the corresponding data set that is closer to the final data set among the corresponding data sets. For example, a competitive situation may include a situation where there are multiple corresponding data sets to which a high-priority policy (e.g., non-recomputation policy) can be applied, but the high-priority policy is applicable only to some of the multiple corresponding data sets.

[0050] For example, the corresponding data set of the fourth partial data set in the first pipeline stage may correspond to the final data set, and there may be competition by priority between the corresponding data set of the fourth partial data set and the corresponding data set of the third partial data set. For example, according to the memory usage status, there may be a situation where the non-recomputation policy can be applied to one of the corresponding data set of the fourth partial data set and the corresponding data set of the third partial data. In this case, the priority can be applied earlier to the training operation of the corresponding data set of the fourth partial data set. Therefore, the fourth partial data set can be trained according to the non-recomputation policy.

[0051] FIG. 8 is a diagram showing an application example of the priority of a storage policy according to an embodiment. Referring to FIG. 8, the first training process 810 shows an example in which a non-recomputation policy is selected for all corresponding data sets of all partial data sets in the pipeline stages 811 to 814. The non-recomputation policy indicates a high memory usage level. For example, the first pipeline stage 811 of the first training process 810 may exceed the memory limit.

[0052] Unlike the first training process 810, the second training process 820 shows an example in which a full recomputation policy is selected for the first corresponding data set of the fourth partial data set in the first pipeline stage 821. In the forward operation of the second training process 820, the activations of the corresponding data sets to which the non-recomputation policy is applied may be stored in memory, and the activations of the corresponding data sets to which the full recomputation policy is applied may also be stored in memory. In the backward operation of the second training process 820, the activations of the corresponding data sets to which the non-recomputation policy is applied are loaded from memory for gradient calculation, and the activations of the corresponding data sets to which the full recomputation policy is applied may be newly generated for gradient calculation.

[0053] Unlike the first training process 810, a full recomputation policy may be arbitrarily applied to the second training process 820. The activations of the corresponding data sets to which the full recomputation policy is applied may be newly generated. Therefore, the memory usage level of the second training process 820 is lower than that of the first training process 820 and does not exceed the memory limit.

[0054] Unlike the second training process 820, the third training process 830 shows an example where a full recomputation policy is selected for the first corresponding data set of the third partial data set in the first pipeline stage 831. When the first corresponding data set of the fourth partial data set and the first corresponding data of the third partial data set compete in terms of priority, the priority may be applied first to the first corresponding data set of the fourth partial data set. Therefore, a non-recomputation policy is applied to the first corresponding data set of the fourth partial data set, and a full recomputation policy is applied to the first corresponding data set of the third partial data set. As a result, the third training process 830 shows a shortened training time compared to the second training process 820 without exceeding the memory limit together with the second training process 820.

[0055] The illustration in FIG. 8 relates to a limited number of pipeline stages 811-814, 821-824, 831-834 and a limited number of partial data sets for the sake of convenience of explanation, but a larger number of pipeline stages and a larger number of partial data sets are used than in reality, and in this case, the training time can be further significantly shortened.

[0056] FIG. 9 is a flowchart exemplarily showing a parallel training method according to an embodiment. Referring to FIG. 9, the electronic device divides a neural network model and a training data set based on pipeline stages for parallel training of the neural network model in step 910, determines a partial model of the neural network model and a partial data set of the training data set, and in step 920, individually selects a storage policy for activations generated by the forward propagation of the partial model for use in calculating the gradient of the backpropagation of the partial model with respect to the corresponding data set of the partial data set processed in the pipeline stage, and in step 930, generates a strategy for parallel training of the neural network model based on the storage policy.

[0057] In the first pipeline stage of the pipeline stage, the first storage policy of the storage policy is used to train the corresponding dataset of the first partial dataset of the partial dataset, and a second storage policy different from the first storage policy can be used to train the corresponding dataset of the second partial dataset of the partial dataset in the first pipeline stage.

[0058] Step 920 includes the step of selecting a storage policy based on the memory usage status.

[0059] The step of selecting a storage policy based on the memory usage status can include the step of selecting a storage policy so that the memory usage is maximized within the available memory capacity.

[0060] The storage policy may include one or more of a full recomputation policy that generates an activation by performing recomputation during backpropagation without storing the activation, a partial recomputation policy that stores some of the activations and performs partial recomputation during backpropagation to generate the remaining part of the activation, and a non-recomputation policy that stores the activation and does not perform recomputation and partial recomputation during backpropagation.

[0061] Step 920 may include the step of selecting a storage policy based on the priority of the storage policy.

[0062] The non-recomputation policy has a higher priority than the partial recomputation policy, and the partial recomputation policy has a higher priority than the full recomputation policy.

[0063] If there is competition by priority among the training operations of the corresponding datasets in the pipeline stage, the priority may be applied earlier to the corresponding dataset closer to the final dataset among the corresponding datasets.

[0064] When a non-recalculation policy is applied to the final data set among the corresponding data sets and one of the other data sets that are not the final data set among the corresponding data sets, the non-recalculation policy can be preferentially applied to the final data set.

[0065] FIG. 10 is a block diagram exemplarily showing the configuration of an electronic device according to an embodiment. Referring to FIG. 10, the electronic device 1000 includes a processor 1010, a memory 1020, a camera 1030, a storage device 1040, an input device 1050, an output device 1060, and a network interface 1070, and these can communicate via a communication bus 1080. For example, the electronic device 1000 may be realized as at least a part of a computing device such as a desktop or a server.

[0066] The processor 1010 executes functions and instructions for execution within the electronic device 1000. For example, the processor 1010 may process instructions stored in the memory 1020 or the storage device 1040. The processor 1010 may perform one or more operations described with reference to FIGS. 1 to 9. The memory 1020 may include a computer-readable storage medium or a computer-readable storage device. The memory 1020 may store instructions for execution by the processor 1010 and store related information while software and / or applications are performed by the electronic device 1000.

[0067] The camera 1030 takes pictures and / or videos. For example, the camera 1030 may take a face image including the user's face. The camera 1030 can provide a 3D image including depth information about an object.

[0068] The storage device 1040 includes a computer-readable storage medium or a computer-readable storage device. The storage device 1040 can store a larger amount of information than the memory 1020 and store the information for a long time. For example, the storage device 1040 may include a magnetic hard disk, an optical disk, a flash memory, a floppy disk, or any other form of non-volatile memory known in this technical field.

[0069] The input device 1050 can receive input from the user through traditional input methods such as a keyboard and a mouse, and new input methods such as touch input, voice input, and image input. For example, the input device 1050 may include a keyboard, a mouse, a touch screen, a microphone, or any other device that can detect input from the user and transmit the detected input to the electronic device 1000.

[0070] The output device 1060 can provide the output of the electronic device 1000 to the user through visual, auditory, or tactile channels. The output device 1060 may include, for example, a display, a touch screen, a speaker, a vibration generator, or any other device that can provide output to the user. The network interface 1070 can communicate with external devices via a wired or wireless network.

[0071] The embodiments described above are implemented by hardware components, software components, or a combination of hardware components and software components. For example, the devices and components described in this embodiment can be implemented using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an ALU (arithmetic logic unit), a digital signal processor, a microcomputer, an FPA (field programmable array), a PLU (programmable logic unit), a microprocessor, or different devices that execute and respond to instructions. The processing device executes an operating system (OS) and one or more software applications executed on the operating system. Further, the processing device accesses, stores, manipulates, processes, and generates data in response to the execution of the software. For the sake of convenience of understanding, the processing device may be described as being one, but those of ordinary skill in the art will understand that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, the processing device may include multiple processors or one processor and one controller. Also, other processing configurations, such as a parallel processor, are possible.

[0072] Software includes a computer program, code, instructions, or a combination of one or more of them, and can configure a processing device to operate as desired or command the processing device independently or in combination. Software and / or data can be permanently or temporarily embodied in any type of machine, component, physical device, virtual device, computer storage medium or device, or signal wave being transmitted. Software can be distributed on a computer system connected to a network and stored or executed in a distributed manner. Software and data can be stored in one or more computer-readable recording media.

[0073] The method according to this embodiment is embodied in the form of program instructions implemented via various computer means and recorded on a computer-readable recording medium. The recording medium includes program instructions, data files, data structures, etc. alone or in combination. The recording medium and program instructions may be specially designed and configured for the purpose of the present invention, or may be known and usable to those skilled in the art of computer software technology. Examples of computer-readable recording media include magnetic media such as hard disks, floppy (registered trademark) disks, and magnetic tapes, optical recording media such as CD-ROMs, DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions such as ROMs, RAMs, flash memories, etc. Examples of program instructions include not only machine language code generated by a compiler but also high-level language code executed by a computer using an interpreter or the like.

[0074] The hardware device described above may be configured to operate as one or more software modules to execute the operations shown in the present invention, and vice versa.

[0075] As described above, the embodiments have been described by way of example with reference to the limited drawings. However, those of ordinary skill in the art can apply various technical modifications and variations based on the above description. For example, the described techniques may be performed in an order different from that described, and / or the components such as the described systems, structures, devices, circuits, etc. may be combined or assembled in a form different from that described, and appropriate results can be achieved even if they are replaced or substituted by other components or equivalents.

[0076] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

Description of Reference Numerals

[0077] 410 First pipeline stage 411 First storage policy 412 Second storage policy 413 Third storage policy

Claims

1. dividing the neural network model and the training dataset to determine a partial model of the neural network model and a partial dataset of the training dataset based on a pipeline stage for parallel training of the neural network model; selecting, individually for a corresponding data set of the partial data set processed in the pipeline stage, a storage policy of activations generated by the forward propagation of the partial model for use in calculating gradients of the back propagation of the partial model; generating a strategy for the parallel training of the neural network model based on the storage policy; A parallel training method comprising:

2. a first one of the pipeline stages is used to train a corresponding dataset of a first one of the partial datasets in a first pipeline stage of the pipeline stages; 2. The parallel training method of claim 1, wherein a second one of the storage policies, different from the first storage policy, is used to train a corresponding dataset of a second partial dataset of the partial dataset in the first pipeline stage.

3. The method of claim 1 , wherein selecting the retention policy comprises selecting the retention policy based on memory usage.

4. The method of claim 3 , wherein selecting the retention policy based on the memory usage state comprises selecting the retention policy such that memory usage is maximized within available memory capacity.

5. The storage policy is: a complete recomputation policy that performs recomputation during the backpropagation to generate the activations without storing the activations; a partial recomputation policy for storing a portion of the activations and performing partial recomputation during the backpropagation to generate a remaining portion of the activations; a no-recomputation policy that stores the activation and does not perform the recomputation and the partial recomputation during the backpropagation; 2. The method of claim 1, further comprising one or more of:

6. The method of claim 5 , wherein selecting the retention policy comprises selecting the retention policy based on a priority of the retention policy.

7. 7. The method of claim 6, wherein the no-recomputation policy has a higher priority than the partial-recomputation policy, and the partial-recomputation policy has a higher priority than the full-recomputation policy.

8. If there is a competition among the training tasks of the corresponding datasets of the pipeline stages according to the priority, the priority is determined by: The parallel training method according to claim 6 , wherein among the corresponding data sets, a corresponding data set closer to a final data set is applied earlier.

9. 7. The parallel training method of claim 6, wherein when the no-recomputation policy is applied to one of a final data set among the corresponding data sets and another data set among the corresponding data sets that is not the final data set, the no-recomputation policy is preferentially applied to the final data set.

10. 1. An electronic device comprising: A processor; A memory for storing instruction words; Including, When the instructions are executed by the processor, the electronic device Based on a pipeline stage for parallel training of a neural network model, dividing the neural network model and the training dataset to determine a partial model of the neural network model and a partial dataset of the training dataset; selecting, individually for the corresponding partial datasets processed in the pipeline stages, a storage policy for the activations generated by the forward propagation of the partial models for use in calculating gradients of the backpropagation of the partial models; An electronic device, wherein a strategy for the parallel training of the neural network model is generated based on the storage policy.

11. a first one of the pipeline stages is used to train a corresponding dataset of a first one of the partial datasets in a first pipeline stage of the pipeline stages; The electronic device of claim 10 , wherein a second storage policy of the storage policies, different from the first storage policy, is used to train a corresponding dataset of a second partial dataset of the partial dataset in the first pipeline stage.

12. The electronic device of claim 10 , wherein the instructions, when executed by the processor, cause the electronic device to select the retention policy based on a memory usage state.

13. The electronic device of claim 10, wherein when the instruction is executed by the processor, the electronic device selects the storage policy based on memory usage conditions, such that the storage policy is selected to maximize memory usage within available memory capacity.

14. The storage policy is: a complete recomputation policy that performs recomputation during the backpropagation to generate the activations without storing the activations; a partial recomputation policy for storing a portion of the activations and performing partial recomputation during the backpropagation to generate a remaining portion of the activations; a no-recomputation policy that stores the activation and does not perform the recomputation and the partial recomputation during the backpropagation; The electronic device of claim 10 , comprising one or more of the following:

15. The electronic device of claim 14 , wherein the instructions, when executed by the processor, cause the electronic device to select the storage policy based on a priority of the storage policies.

16. 16. The electronic device of claim 15, wherein the no-recomputation policy has a higher priority than the partial-recomputation policy, and the partial-recomputation policy has a higher priority than the full-recomputation policy.

17. If there is a competition among the training tasks of the corresponding datasets of the pipeline stages according to the priority, the priority is determined by: The electronic device according to claim 15 , wherein among the corresponding data sets, a corresponding data set closer to a final data set is applied earlier.

18. 1. A method for training a neural network model, comprising: calculating a first value in a first pipeline stage of the neural network model based on a first partial data set; processing the first value according to a first storage policy based on the first pipeline stage and the first partial data set; calculating a second value in the first pipeline stage of the neural network model based on a second partial data set; processing the second value according to a second storage policy based on the first pipeline stage and the second partial data set; calculating a third value in a second pipeline stage of the neural network model based on the first partial data set; processing the third value according to a third storage policy based on the second pipeline stage and the first partial data set; calculating a fourth value in the second pipeline stage of the neural network model based on the second partial data set; processing the fourth value according to a fourth storage policy based on the second pipeline stage and the second partial data set; A training method including:

19. 20. The method of claim 18, wherein the third value and the second value are calculated simultaneously.

20. The first storage policy, the second storage policy, the third storage policy, and the fourth storage policy are a full recomputation policy that performs recomputation during backpropagation to generate the value without storing the value; a partial recalculation policy that stores a portion of the value and performs a partial recalculation during the backpropagation to generate a remaining portion of the value; a no-recomputation policy that stores the value and does not perform the recomputation and the partial recomputation during the backpropagation; 20. The method of claim 18, wherein the storage policy type is selected from a set of storage policy types including any one or more of: