Training device, method and related equipment for a neural network model

By storing the weight coefficients of the neural network model and the initial variables of the optimizer, and using the set communication to aggregate the weight coefficients and gradients, the problem of insufficient memory of a single block accelerator is solved, and efficient training of large neural network models is achieved.

CN113705801BActive Publication Date: 2025-06-03HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010441573.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-05-22
Publication Date
2025-06-03
Estimated Expiration
2040-05-22

AI Technical Summary

Technical Problem

When training larger neural network models, the single block accelerator lacks video memory space, resulting in the inability to train.

Method used

By storing the complete weight coefficients in the neural network model and the complete initial variables of the optimizer in multiple accelerators in distributed storage and converging them through collective communication to obtain the complete weight coefficients and target gradients, the memory consumption of each accelerator is reduced.

Benefits of technology

It effectively reduces the memory consumption of the training device during the neural network model training process, avoids the problem of insufficient memory, and improves the ability to train large neural network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113705801B_ABST
    Figure CN113705801B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a training device, method, and related equipment for a neural network model. The training device can be used in the scenario of training a neural network in the field of artificial intelligence (AI). The training device includes multiple accelerators. During the parallel processing of training a neural network model using multiple accelerators in the training device, the complete weight coefficients in the neural network model are distributed and stored in the multiple accelerators in the training device. Subsequently, the complete weight coefficients are obtained through aggregation in the multiple accelerators, and then the neural network model is further trained on each accelerator according to different input data and the complete weight coefficients. That is, by distributing and storing the complete weight coefficients in the multiple accelerators in the training device, the video memory consumption of the training device during the training process of the neural network model can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a training device, method, and related equipment for a neural network model. Background Art

[0002] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. The research in the field of artificial intelligence includes robots, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, AI basic theory, etc.

[0003] In recent years, the training of neural network models has been developing towards large networks and large amounts of data. Generally speaking, the computing requirements that surge in the training network can be improved through data parallelism. The basic idea of data parallelism is to use model replicas on multiple devices to train data subsets simultaneously and synchronize the model parameters across replicas at the end of each iteration.

[0004] Specifically, Figure 1A schematic diagram of training a neural network using data parallelism is given. This training process is implemented by a CPU and multiple accelerators (Accelerator 1, Accelerator 2, ..., Accelerator n). Among them, multiple accelerators train together. Generally, the training process includes the following steps: 1) Create the same training model on each accelerator. For example, when training a Bert network, each accelerator needs to have a complete Bert model; 2) Initialize the weight coefficients on Accelerator 1 and send these weight coefficients to each accelerator through the broadcast operation in collective communication (1001). Generally, when training a neural network from scratch, a random method can be used to assign an initial value to the weights. To make the initial weight values on each accelerator consistent, the method of randomly initializing the weights on any one accelerator first and then sending these weights to each accelerator is adopted; 3) The CPU sends different data to different accelerators; 4) Each accelerator performs forward and backward calculations to obtain the corresponding gradient information. This step is an operation inside each accelerator. After forward and backward calculations, the gradient information corresponding to the batch data of the current accelerator is obtained. Since the input data is ensured to be different in 3), the gradients obtained by each accelerator are different; 5) Perform the allreduce operation in collective communication to obtain the average gradient. Average the different gradients obtained by each accelerator in 4). After this step, the gradient values on all accelerators will be consistent, all being the average value of the gradients of each accelerator; 6) Use the average gradient value to update the initial weights. Since the operation in 2) ensures that the initial weights on each accelerator are consistent, and the update amount on each accelerator is the average gradient after being averaged by allreduce. Therefore, it can be ensured that the weight values on each accelerator can also be kept consistent after each update. Among them, in the update process of 6), each accelerator can further use this average gradient value and the weights obtained in 1) as the input of the optimizer, perform optimization operations through the initial variables (1002) of the optimizer, the optimizer outputs the processed gradient, and each accelerator further uses the processed gradient to update the weights.

[0005] Among them, the above accelerators can be a graphics processing unit (GPU), a neural processing unit (NPU), or a tensor processing unit (TPU); gradient aggregation can be implemented by various methods, such as collective communication.

[0006] In the above data parallel processing method, training parameters such as the initial weights in step 2) and the initial variables in step 6) consume the storage space used in accelerator training, such as the video memory used in GPU training and the memory used in CPU training. When training a relatively large neural network model, there will be a problem that the storage space in a single accelerator is insufficient, resulting in the inability to perform training. Summary of the Invention

[0007] The embodiments of the present application provide a training device, method and related equipment for a neural network model, which are used to reduce the video memory consumption of the training device during the training process of the neural network model.

[0008] In a first aspect of the embodiments of the present application, a training device for a neural network model is provided. The training device includes multiple accelerators. During the process of the training device training the neural network model, each accelerator in the training device is used to store a part of the weight coefficients. The part of the weight coefficients stored by each of the multiple accelerators forms a complete weight coefficient, that is, the complete weight coefficient in the neural network model is distributed and stored in multiple accelerators in the training device. Then, each accelerator converges the part of the weight coefficients stored separately in the multiple accelerators to obtain the complete weight coefficient. After that, each accelerator trains the neural network model according to the input data and the complete weight coefficient, where the input data of the multiple accelerators are different from each other. During the parallel processing of using multiple accelerators by the training device to train the neural network model, the complete weight coefficient in the neural network model is distributed and stored in multiple accelerators in the training device, and then the complete weight coefficient is obtained through convergence in the multiple accelerators. On each accelerator, the neural network model is further trained according to different input data and the complete weight coefficient, that is, by distributing and storing the complete weight coefficient in multiple accelerators in the training device, the video memory consumption of the training device during the training process of the neural network model is reduced.

[0009] In a possible implementation manner of the first aspect of the embodiments of the present application, when each accelerator trains the neural network model according to the input data and the complete weight coefficients, it is specifically used to calculate gradient information according to the input data and the complete weight coefficients; each accelerator then calculates the target gradient according to the gradient information of the multiple accelerators; further, each accelerator uses the target gradient to update the partial weight coefficients and trains the neural network model according to the updated partial weight coefficients. Among them, after each accelerator in the training device calculates the gradient information according to different input data and complete weight coefficients, it then calculates the target gradient for updating the partial weight coefficients according to the gradient information of the multiple accelerators. Further, each accelerator uses the target gradient to update the partial weight coefficients stored in each accelerator and trains the neural network model according to the updated partial weight coefficients. This implementation manner provides a specific implementation process for training the neural network model according to different input data, improves the feasibility of the solution, and thus improves the implementation flexibility of the present solution.

[0010] In a possible implementation manner of the first aspect of the embodiments of the present application, each accelerator in the training device is further used to store some initial variables in the optimizer, and the partial initial variables stored by the multiple accelerators respectively form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model; among them, when each accelerator in the training device uses the target gradient to update the partial weight coefficients, it is specifically used to process the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient; thereafter, each accelerator updates the partial weight coefficients according to the processed target gradient. Among them, when using the optimizer to optimize the weight coefficients of the neural network model, each accelerator can perform distributed storage on the initial variables of the accelerator, that is, each accelerator stores some initial variables, and each accelerator then uses the target gradient and the partial weight coefficients as the input of the optimizer, and performs optimization processing through the preset optimization algorithm in the optimizer to obtain the processed target gradient, and then updates the partial weight coefficients stored in each accelerator according to the processed target gradient, that is, by distributedly storing the complete initial variables in the optimizer in the multiple accelerators in the training device, thereby further reducing the video memory consumption of the training device during the training process of the neural network model.

[0011] In a possible implementation of the first aspect of the embodiments of the present application, the optimizer includes vector operations. When each accelerator in the training device processes the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient, it is specifically used to calculate the scalar representation of the target gradient; then converge the scalar representations of the target gradient in the multiple accelerators to obtain the summation result of the target gradient; thereafter, each accelerator calculates the vector representation of the target gradient according to the summation result; and further processes the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient. Among them, if the optimizer includes appropriate operations (such as matrix operations, vector operations, or other vector operations, etc.), complete gradients are required for calculation. Since the target gradient is distributed on each accelerator, when each accelerator uses the accelerator to obtain the processed target gradient, each accelerator first calculates the scalar representation of the target gradient; then converges the scalar representations of the target gradient in the multiple accelerators to obtain the summation result of the target gradient; thereafter, each accelerator calculates the vector representation of the target gradient according to the summation result; and further processes the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient. In this implementation, the solution can be applied to the implementation process of an optimizer containing vector operations, thereby improving the feasibility of the solution.

[0012] In a possible implementation of the first aspect of the embodiments of the present application, when each accelerator in the training device converges the scalar representations of the target gradient in the multiple accelerators to obtain the summation result of the target gradient, it is specifically used to converge the scalar representations of the target gradient in the multiple accelerators through the reduction operation (allreduce) in the collective communication method to obtain the summation result of the target gradient. Among them, each accelerator in the training device can obtain the summation result of the target gradient through the reduction operation (allreduce) in the collective communication method among the multiple accelerators. This implementation provides a specific implementation process for obtaining the summation result of the target gradient, improves the feasibility of the solution, and thus improves the implementation flexibility of the present solution.

[0013] In a possible implementation manner of the first aspect of the embodiments of the present application, the partial weight coefficients include the weight coefficients obtained by evenly dividing the complete weight coefficients and distributing them to the multiple accelerators one by one. During the training process of the neural network model, the processing capabilities of the multiple accelerators in the training device are generally the same or nearly the same. Therefore, the complete weight coefficients can be evenly divided according to the number of the multiple accelerators and distributed to the multiple accelerators one by one, so that each accelerator stores the evenly divided partial weight coefficients in a distributed manner. In this implementation manner, a specific implementation process for the distributed storage of the complete weight coefficients in the multiple accelerators is provided, improving the feasibility of the solution, and thus enhancing the implementation flexibility of the present solution.

[0014] In a possible implementation manner of the first aspect of the embodiments of the present application, when each accelerator in the training device aggregates the partial weight coefficients respectively stored in the multiple accelerators to obtain the complete weight coefficient, it is specifically configured to aggregate the partial weight coefficients respectively stored in the multiple accelerators through the Allgather operation in the collective communication mode to obtain the complete weight coefficient. Each accelerator in the training device can obtain the complete weight coefficient through the Allgather operation in the collective communication mode among the multiple accelerators. This implementation manner provides a specific implementation process for obtaining the complete weight coefficient, improving the feasibility of the solution, and thus enhancing the implementation flexibility of the present solution.

[0015] In a possible implementation manner of the first aspect of the embodiments of the present application, when each accelerator in the training device calculates the target gradient according to the gradient information of the multiple accelerators, it is specifically configured to calculate the target gradient according to the gradient information of the multiple accelerators through the ReduceScatter operation in the collective communication mode. Each accelerator in the training device can calculate and obtain the target gradient through the ReduceScatter operation in the collective communication mode among the multiple accelerators. This implementation manner provides a specific implementation process for calculating and obtaining the target gradient, improving the feasibility of the solution, and thus enhancing the implementation flexibility of the present solution.

[0016] In a possible implementation manner of the first aspect of the embodiments of the present application, the partial initial variables include the initial variables obtained by evenly dividing the complete initial variables and distributing them to the multiple accelerators one by one. During the training process of the neural network model, the processing capabilities of the multiple accelerators in the training device are generally the same or nearly the same. Therefore, the complete initial variables can be evenly divided according to the number of the multiple accelerators and distributed to the multiple accelerators one by one, so that each accelerator stores the evenly divided partial initial variables in a distributed manner. In this implementation manner, a specific implementation process for the distributed storage of the complete initial variables in the multiple accelerators is provided, improving the feasibility of the solution, and thus enhancing the implementation flexibility of the present solution.

[0017] In a possible implementation manner of the first aspect of the embodiments of the present application, each accelerator in the training device is further configured to: obtain update parameters of this part of the weight coefficients, and update this part of the weight coefficients according to the update parameters of this part of the weight coefficients; and / or, each accelerator is further configured to obtain update parameters of the initial variable, and update the initial variable according to the update parameters of the variable; and / or, each accelerator is further configured to obtain update parameters of the target gradient, and update the target gradient according to the update parameters of the target gradient; and / or, each accelerator is further configured to obtain update parameters of the processed target gradient, and update the target gradient according to the update parameters of the processed target gradient. Wherein, the target parameters involved in the training process of the neural network model can be stored distributively, and the target parameters include this part of the weight coefficients and / or the initial variable and / or the target gradient and / or the processed target gradient, etc. Thus, when there is an update to this part of the weight coefficients corresponding to the neural network model or the initial variable in the optimizer, the distributively stored target parameters can be updated respectively in each accelerator in the training device, thereby further reducing the video memory consumption of the accelerators in the training device.

[0018] The second aspect of the embodiments of the present application provides a training device for a neural network model. The training device includes a plurality of accelerators. During the process of training the neural network model by the training device, each accelerator in the training device is configured to: calculate gradient information according to input data and complete weight coefficients, where the input data of the plurality of accelerators are different from each other; then, each accelerator calculates a target gradient according to the gradient information of the plurality of accelerators; each accelerator is further configured to store some initial variables in an optimizer, and the some initial variables stored by the plurality of accelerators respectively form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model; thereafter, each accelerator processes the target gradient and some weight coefficients according to the some initial variables to obtain a processed target gradient, and the some weight coefficients processed by the plurality of accelerators respectively form the complete weight coefficients; each accelerator further updates the complete weight coefficients according to the processed target gradient, and trains the neural network model according to the updated complete weight coefficients. Among them, during the parallel processing of training the neural network model by using a plurality of accelerators in the training device, the complete initial variables of the optimizer in the neural network model are distributed and stored in the plurality of accelerators in the training device. Each accelerator then processes the target gradient and some weight coefficients according to the some initial variables to obtain a processed target gradient. Thereafter, each accelerator further updates the complete weight coefficients according to the processed target gradient, and trains the neural network model according to the updated complete weight coefficients, that is, by distributing and storing the complete initial weights of the optimizer in the plurality of accelerators in the training device, thereby reducing the video memory consumption of the training device during the training process of the neural network model.

[0019] In a possible implementation manner of the second aspect of the embodiments of the present application, the optimizer includes vector operations. When each accelerator in the training device processes the target gradient and partial weight coefficients according to the partial initial variables to obtain the processed target gradient, it is specifically used to calculate the scalar representation of the target gradient; each accelerator then converges the scalar representations of the target gradient in the multiple accelerators to obtain the summation result of the target gradient; then, each accelerator calculates the vector representation of the target gradient according to the summation result; thereafter, each accelerator processes the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient. Among them, if the optimizer includes appropriate operations (such as matrix operations, vector operations, or other vector operations, etc.), complete gradients are required to participate in the calculation. Since the target gradient is distributed on each accelerator, when each accelerator uses the accelerator to obtain the processed target gradient, each accelerator first calculates the scalar representation of the target gradient; then converges the scalar representations of the target gradient in the multiple accelerators to obtain the summation result of the target gradient; thereafter, each accelerator calculates the vector representation of the target gradient according to the summation result; and further processes the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient. In this implementation manner, the solution can be applied to the implementation process of an optimizer including vector operations, thereby improving the feasibility of the solution.

[0020] In a possible implementation manner of the second aspect of the embodiments of the present application, when each accelerator in the training device converges the scalar representations of the target gradient in the multiple accelerators to obtain the summation result of the target gradient, it is specifically used to converge the scalar representations of the target gradient in the multiple accelerators through a reduction operation in the collective communication method to obtain the summation result of the target gradient. Among them, each accelerator in the training device can obtain the summation result of the target gradient through a reduction operation (allreduce) in the collective communication method among multiple accelerators. This implementation manner provides a specific implementation process for obtaining the summation result of the target gradient, improves the feasibility of the solution, and thus improves the implementation flexibility of the present solution.

[0021] In a possible implementation manner of the second aspect of the embodiments of the present application, when each accelerator calculates the target gradient according to the gradient information of the multiple accelerators, it is specifically configured to calculate the target gradient through a ReduceScatter operation in the collective communication method according to the gradient information of the multiple accelerators. Each accelerator in the training device can calculate the target gradient through the ReduceScatter operation in the collective communication method among the multiple accelerators. This implementation manner describes the specific implementation process of calculating the target gradient, improves the feasibility of the solution, and thus enhances the implementation flexibility of the present solution.

[0022] In a possible implementation manner of the second aspect of the embodiments of the present application, the partial initial variables include the initial variables obtained by evenly dividing the complete initial variables and distributing them to the multiple accelerators one by one. During the training process of the neural network model, the processing capabilities of the multiple accelerators in the training device are generally the same or nearly the same. Therefore, the complete initial variables can be evenly divided according to the number of the multiple accelerators and distributed to the multiple accelerators one by one, so that each accelerator stores the evenly divided partial initial variables in a distributed manner. In this implementation manner, the specific implementation process of the distributed storage of the complete initial variables among the multiple accelerators is provided, which improves the feasibility of the solution and thus enhances the implementation flexibility of the present solution.

[0023] In a possible implementation manner of the second aspect of the embodiments of the present application, each accelerator in the training device is further configured to: obtain the update parameter of the complete weight coefficient and update the complete weight coefficient according to the update parameter of the complete weight coefficient; and / or, each accelerator is further configured to obtain the update parameter of the initial variable and update the initial variable according to the update parameter of the initial variable; and / or, each accelerator is further configured to obtain the update parameter of the target gradient and update the target gradient according to the update parameter of the target gradient; and / or, each accelerator is further configured to obtain the update parameter of the processed target gradient and update the target gradient according to the update parameter of the processed target gradient. The target parameters involved in the training process of the neural network model can be stored in a distributed manner. The target parameters include the complete weight coefficient and / or the initial variable and / or the target gradient and / or the processed target gradient, etc. Thus, when there is an update to the corresponding partial weight coefficient of the neural network model or the initial variable in the optimizer, the target parameters stored in a distributed manner can be updated respectively in each accelerator in the training device, thereby further reducing the video memory consumption of the accelerators in the training device.

[0024] A third aspect of the embodiments of the present application provides a method for training a neural network model. The training method is applied to multiple accelerators, and the multiple accelerators are included in a training device. The method includes: storing partial weight coefficients, and the partial weight coefficients stored by each of the multiple accelerators form a complete weight coefficient; aggregating the partial weight coefficients separately stored in the multiple accelerators to obtain the complete weight coefficient; training the neural network model according to input data and the complete weight coefficient, where the input data of the multiple accelerators are different from each other.

[0025] In a possible implementation manner of the third aspect of the embodiments of the present application, the training the neural network model according to the input data and the complete weight coefficient includes: calculating gradient information according to the input data and the complete weight coefficient; calculating a target gradient according to the gradient information of the multiple accelerators; using the target gradient to update the partial weight coefficients, and training the neural network model according to the updated partial weight coefficients.

[0026] In a possible implementation manner of the third aspect of the embodiments of the present application, the method further includes: storing partial initial variables in an optimizer, and the partial initial variables stored by each of the multiple accelerators form the complete initial variables of the optimizer, where the optimizer is used to update the weight coefficients of the neural network model; the using the target gradient to update the partial weight coefficients includes: processing the target gradient and the partial weight coefficients according to the partial initial variables to obtain a processed target gradient; updating the partial weight coefficients according to the processed target gradient.

[0027] In a possible implementation manner of the third aspect of the embodiments of the present application, the optimizer includes vector operations, and the processing the target gradient and the partial weight coefficients according to the partial initial variables to obtain a processed target gradient includes: calculating a scalar representation of the target gradient; aggregating the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient; calculating a vector representation of the target gradient according to the summation result; processing the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient.

[0028] In a possible implementation manner of the third aspect of the embodiments of the present application, the aggregating the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient includes: aggregating the scalar representations of the target gradient in the multiple accelerators through a reduction operation in a collective communication manner to obtain a summation result of the target gradient.

[0029] In a possible implementation manner of the third aspect of the embodiments of the present application, the partial weight coefficients include weight coefficients obtained by evenly dividing the complete weight coefficient and distributing them to the multiple accelerators one by one.

[0030] In a possible implementation manner of the third aspect of the embodiments of the present application, aggregating the partial weight coefficients respectively stored in the multiple accelerators to obtain the complete weight coefficient includes: aggregating the partial weight coefficients respectively stored in the multiple accelerators through a gather operation in the collective communication manner to obtain the complete weight coefficient.

[0031] In a possible implementation manner of the third aspect of the embodiments of the present application, calculating the target gradient according to the gradient information of the multiple accelerators includes: calculating the target gradient according to the gradient information of the multiple accelerators through a reduce scatter operation in the collective communication manner.

[0032] In a possible implementation manner of the third aspect of the embodiments of the present application, the partial initial variables include the initial variables obtained by evenly dividing the complete initial variables and respectively allocating them to the multiple accelerators.

[0033] In a possible implementation manner of the third aspect of the embodiments of the present application, the method further includes: obtaining an update parameter of the partial weight coefficients, and updating the partial weight coefficients according to the update parameter of the partial weight coefficients; and / or, obtaining an update parameter of the initial variables, and updating the initial variables according to the update parameter of the initial variables; and / or, obtaining an update parameter of the target gradient, and updating the target gradient according to the update parameter of the target gradient; and / or, obtaining an update parameter of the processed target gradient, and updating the target gradient according to the update parameter of the processed target gradient.

[0034] For the specific implementation steps of the third aspect of the present application and various possible implementation manners of the third aspect, as well as the beneficial effects brought by each possible implementation manner, reference may be made to the descriptions in various possible implementation manners of the first aspect, and details are not described herein again.

[0035] The fourth aspect of the embodiments of the present application provides a method for training a neural network model. The training method is applied to multiple accelerators, and the multiple accelerators are included in a training device. The method includes: calculating gradient information according to input data and complete weight coefficients, where the input data of the multiple accelerators are different from each other; calculating a target gradient according to the gradient information of the multiple accelerators; storing partial initial variables in an optimizer, and the partial initial variables respectively stored by the multiple accelerators form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model; processing the target gradient and partial weight coefficients according to the partial initial variables to obtain a processed target gradient, and the partial weight coefficients respectively processed by the multiple accelerators form the complete weight coefficient; updating the complete weight coefficient according to the processed target gradient, and training the neural network model according to the updated complete weight coefficient.

[0036] In a possible implementation of the fourth aspect of the embodiments of the present application, the optimizer includes vector operations. The processing of the target gradient and partial weight coefficients according to the partial initial variables to obtain the processed target gradient includes: calculating a scalar representation of the target gradient; aggregating the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient; calculating a vector representation of the target gradient according to the summation result; and processing the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient.

[0037] In a possible implementation of the fourth aspect of the embodiments of the present application, the aggregating the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient includes: aggregating the scalar representations of the target gradient in the multiple accelerators through a reduction operation in the collective communication method to obtain a summation result of the target gradient.

[0038] In a possible implementation of the fourth aspect of the embodiments of the present application, the calculating the target gradient according to the gradient information of the multiple accelerators includes: calculating the target gradient through a reduce-scatter operation in the collective communication method according to the gradient information of the multiple accelerators.

[0039] In a possible implementation of the fourth aspect of the embodiments of the present application, the partial initial variables include the initial variables obtained by evenly dividing the complete initial variables and distributing them to the multiple accelerators one by one.

[0040] In a possible implementation of the fourth aspect of the embodiments of the present application, the method further includes: obtaining an update parameter of the complete weight coefficient and updating the complete weight coefficient according to the update parameter of the complete weight coefficient; and / or, obtaining an update parameter of the initial variable and updating the initial variable according to the update parameter of the initial variable; and / or, obtaining an update parameter of the target gradient and updating the target gradient according to the update parameter of the target gradient; and / or, obtaining an update parameter of the processed target gradient and updating the target gradient according to the update parameter of the processed target gradient.

[0041] For the specific implementation steps of the fourth aspect of the present application and various possible implementation manners of the fourth aspect, as well as the beneficial effects brought by each possible implementation manner, reference may be made to the descriptions in various possible implementation manners in the second aspect, which will not be elaborated herein one by one.

[0042] The fifth aspect of the embodiments of the present application provides a training device for a neural network model. The training device includes a plurality of accelerators, and each accelerator includes: a storage unit for storing partial weight coefficients, and the partial weight coefficients stored by each of the plurality of accelerators form a complete weight coefficient; an aggregation unit for aggregating the partial weight coefficients separately stored in the plurality of accelerators to obtain the complete weight coefficient; and a training unit for training the neural network model according to input data and the complete weight coefficient, wherein the input data of the plurality of accelerators are different from each other.

[0043] In the fifth aspect of the present application, the constituent modules of the training device of the neural network can also be used to execute the steps executed by the training device in each possible implementation manner of the first aspect. Specifically, reference can be made to the first aspect, and details are not described herein again.

[0044] The sixth aspect of the embodiments of the present application provides a training device for a neural network model. The training device includes a plurality of accelerators, and each accelerator includes: a calculation unit for calculating gradient information according to input data and a complete weight coefficient, wherein the input data of the plurality of accelerators are different from each other; the calculation unit is further configured to calculate a target gradient according to the gradient information of the plurality of accelerators; a storage unit for storing partial initial variables in an optimizer, and the partial initial variables stored by each of the plurality of accelerators form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model; a processing unit for processing the target gradient and partial weight coefficients according to the partial initial variables to obtain a processed target gradient, and the partial weight coefficients processed by each of the plurality of accelerators form the complete weight coefficient; and an update unit for updating the complete weight coefficient according to the processed target gradient and training the neural network model according to the updated complete weight coefficient.

[0045] In the sixth aspect of the present application, the constituent modules of the acquisition device of the neural network can also be used to execute the steps executed by the training device in each possible implementation manner of the second aspect. Specifically, reference can be made to the second aspect, and details are not described herein again.

[0046] The seventh aspect of the embodiments of the present application provides a computer-readable storage medium. The computer storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method of the third aspect or any possible implementation manner of the third aspect described above, or when the program instructions are executed by a processor, the processor is caused to execute the method of the fourth aspect or any possible implementation manner of the fourth aspect described above.

[0047] The eighth aspect of the embodiments of the present application provides a chip system, which includes a processor for supporting an access network device to implement the functions involved in the above-mentioned third aspect or any possible implementation manner of the third aspect, and the above-mentioned fourth aspect or any possible implementation manner of the fourth aspect. In a possible design, the chip system may further include a memory for storing necessary program instructions and data of the access network device. The chip system may be composed of chips or may include chips and other discrete devices.

[0048] The ninth aspect of the embodiments of the present application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor executes the method described in the above-mentioned third aspect or any possible implementation manner of the third aspect, or the processor executes the method described in the above-mentioned fourth aspect or any possible implementation manner of the fourth aspect.

[0049] Among them, the technical effects brought by the third to ninth aspects or any possible implementation manner thereof can be referred to the technical effects brought by the first aspect or different possible implementation manners of the first aspect, or refer to the technical effects brought by the second aspect or different possible implementation manners of the second aspect, which will not be elaborated here.

[0050] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages: During the parallel processing of training a neural network model by a training device using multiple accelerators, the complete weight coefficients in the neural network model are distributed and stored in multiple accelerators in the training device, and then the complete weight coefficients are obtained by aggregation in the multiple accelerators. Further, on each accelerator, the neural network model is trained according to different input data and the complete weight coefficients, that is, the complete weight coefficients are distributed and stored in multiple accelerators in the training device, thereby reducing the video memory consumption of the training device during the training process of the neural network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0052] Figure 1 It is a schematic diagram of the training process of the neural network model provided by the embodiments of the present application;

[0053] Figure 2 It is another schematic diagram of the training process of the neural network model provided by the embodiments of the present application;

[0054] Figure 3 Another schematic diagram of the neural network model training process provided by the embodiments of this application;

[0055] Figure 4 A schematic diagram of the realization of collective communication in the neural network model training process provided by the embodiments of this application;

[0056] Figure 5 Another schematic diagram of the realization of collective communication in the neural network model training process provided by the embodiments of this application;

[0057] Figure 6 Another schematic diagram of the realization of collective communication in the neural network model training process provided by the embodiments of this application;

[0058] Figure 7 Another schematic diagram of the realization of collective communication in the neural network model training process provided by the embodiments of this application;

[0059] Figure 8 A schematic diagram of the system architecture of the neural network model training process provided by the embodiments of this application;

[0060] Figure 9 Another schematic diagram of the system architecture of the neural network model training process provided by the embodiments of this application;

[0061] Figure 10 Another schematic diagram of the neural network model training process provided by the embodiments of this application;

[0062] Figure 11 A schematic diagram of the training device implementing the neural network model training method provided by the embodiments of this application;

[0063] Figure 12 Another schematic diagram of the training device implementing the neural network model training method provided by the embodiments of this application;

[0064] Figure 13 Another schematic diagram of the training device implementing the neural network model training method provided by the embodiments of this application;

[0065] Figure 14 Another schematic diagram of the training device implementing the neural network model training method provided by the embodiments of this application;

[0066] Figure 15 A schematic diagram of the training device provided by the embodiments of this application;

[0067] Figure 16 Another schematic diagram of the training device provided by the embodiments of this application;

[0068] Figure 17 Another schematic diagram of the training device provided by the embodiment of the present application. Detailed implementation manners

[0069] Next, the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention. Terms such as "first", "second", "third", and "fourth" in the specification and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. The mention of "embodiment" in this article means that a specific feature, structure, or characteristic described in conjunction with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0070] First, some terms in the present application are explained to facilitate the understanding of those skilled in the art.

[0071] 1. Neural network training

[0072] The basic structure of the neural network is as Figure 2 shown. The input x undergoes multiple transformations to obtain the output y. Generally speaking, the multiple transformations include two major categories. One is the linear layer (f(x, w) in the figure), generally a convolutional layer, a fully connected layer, etc. This type of layer often has learnable weight coefficients. The other is the non-linear layer (a(x) in the figure), also called the activation layer, generally a ReLU layer, a sigmoid layer, etc. The linear layer and the non-linear layer are alternately stacked and act layer by layer to finally obtain the output.

[0073] Mathematically, the neural network can be regarded as a layer-by-layer function transformation. For example, Figure 2 in, that is, it describes a function of y = a(f(x, w)). In supervised training, the output y needs to fit the artificially labeled label. By adjusting the weight coefficient w, the neural network can fit the relationship between different y and x. The process of continuously adjusting the weight coefficient w to fit the relationship between x and the label is called neural network training.

[0074] The schematic diagram of neural network training is as Figure 3As shown. Neural network training generally uses the backpropagation method for calculation. The purpose is to calculate the gradient dw of the weight coefficient in reverse through the difference between the forward predicted value y and the calibrated standard value label, and then adjust the value of the weight through the gradient to make the error between y and label smaller. Such a process of forward, backward, and weight adjustment is called one iteration. The neural network continuously adjusts the parameters by repeating such iterations from tens of thousands to millions of times, and finally obtains a better solution.

[0075] After obtaining the gradient dw, there are many adjustment strategies, that is, optimizers. Basic ones such as stochastic gradient descent simply multiply dw by the learning rate and then adjust the weights. Currently, widely used ones such as the Adam optimizer and the Lar optimizer, etc.

[0076] 2. Collective Communication

[0077] Collective communication provides application programming interfaces (APIs) for many operations for users to use, to complete operations such as averaging in multiple accelerators. Correspondingly, in data parallel neural network training, the averaging of gradients on each accelerator can be completed through collective communication operations. Several basic collective communication operations are as Figures 4 to 7 shown. Taking multiple accelerators including rank0, rank1, rank2, and rank3 as an example, specifically:

[0078] Allreduce: Sum the data on each accelerator. After the allreduce operation, the same summation result is obtained on each accelerator. The Allreduce process is as Figure 4 shown.

[0079] Broadcast: Copy the data on a certain accelerator to all accelerators. The Broadcast process is as Figure 5 shown.

[0080] Allgather: Concatenate the contents on each accelerator, and each accelerator can obtain the synthesized large tensor. The Allgather process is as Figure 6 shown.

[0081] ReduceScatter: After splitting the tensor on each accelerator, each accelerator gets the corresponding different part. The ReduceScatter process is as Figure 7 shown.

[0082] See Appendix Figure 8, an embodiment of the present application provides a system architecture 200 for the neural network training process. The data acquisition device 260 is used to acquire data and store it in the database 230. The training device 220 generates a target model / rule 201 based on the text data maintained in the database 230. Exemplarily, when the target model / rule is used for training natural language processing (NLP), the data involved in this system can be text data. Below, taking the training process of NLP as an example, it will be described in more detail how the training device 220 obtains the target model / rule 201 based on the text data.

[0083] The operation of each layer in a deep neural network can be described by a mathematical expression as follows: From a physical perspective, the operation of each layer in a deep neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of the matrix) through five operations on the input space. These five operations include: 1. Dimension increase / dimension decrease; 2. Enlargement / shrinkage; 3. Rotation; 4. Translation; 5. "Bending". Among them, the operations of 1, 2, and 3 are completed by , the operation of 4 is completed by +b, and the operation of 5 is implemented by a(). The reason for using the word "space" here is that the object to be classified is not a single thing, but a class of things. Space refers to the set of all individuals of this class of things. Among them, W is the weight vector, and each value in this vector represents the weight value of a neuron in this layer of the neural network. This vector W determines the space transformation from the input space to the output space described above, that is, the weight W of each layer controls how to transform the space. The purpose of training a deep neural network, that is, ultimately obtaining the weight matrix of all layers of the trained neural network (the weight matrix formed by many layers of vectors W). Therefore, the training process of a neural network is essentially to learn the way to control space transformation, and more specifically, to learn the weight matrix.

[0084] Since we hope that the output of the deep neural network is as close as possible to the value we truly want to predict, we can compare the predicted value of the current network with the true target value, and then update the weight vector of each layer of the neural network according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network.) For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the neural network can predict the true target value. Therefore, we need to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0085] The target model / rule obtained by the training device 220 can be applied to different systems or devices. In the appendix Figure 8 The execution device 210 is configured with an I / O interface 212 to interact with external devices. The "user" can input data to the I / O interface 212 through the client device 240.

[0086] The execution device 210 can call data, code, etc. in the data storage system 250, and can also store data, instructions, etc. in the data storage system 250.

[0087] The calculation module 211 processes the input data using the target model / rule 201 to obtain gradient information, and then further optimizes the gradient information through the associated function module 213 to obtain a processing result. Among them, the associated function module 213 can specifically include an optimizer, that is, use the preset optimization algorithm in the optimizer to optimize the gradient information.

[0088] Finally, the I / O interface 212 returns the processing result to the client device 240 and provides it to the user.

[0089] More deeply, the training device 220 can generate corresponding target models / rules 201 based on different data for different targets to provide better results for users.

[0090] In the appendix Figure 8In the case shown, the user can manually specify the data in the input execution device 210. For example, the user can operate in the interface provided by the I / O interface 212. In another case, the client device 240 can automatically input data to the I / O interface 212 and obtain the result. If the client device 240 needs the user's authorization to automatically input data, the user can set the corresponding permissions in the client device 240. The user can view the result output by the execution device 210 on the client device 240, and the specific presentation form can be display, sound, action, etc. The client device 240 can also be used as a data collection end to store the collected text data in the database 230.

[0091] It should be noted that the appendix Figure 8 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in the appendix Figure 8 the data storage system 250 is an external memory relative to the execution device 210. In other cases, the data storage system 250 can also be placed in the execution device 210.

[0092] In Figure 8 the shown system architecture, generally speaking, a single accelerator is used in the training device 220 to implement the neural network model. The accelerator can be a GPU, TPU, NPU, or other accelerators, etc. In recent years, the neural network has developed in the direction of large networks and large amounts of data, resulting in a surge in computing requirements. The single accelerator in the common training device can no longer meet the needs of training the network. Therefore, a very large number of parallel methods have emerged, such as data parallelism, model parallelism, hybrid parallelism, pipeline parallelism, etc. Among them, the most common is data parallelism. The following will introduce Figure 9 the system architecture for implementing the training process of the neural network model through data parallelism.

[0093] As Figure 9 shown, in the system architecture implemented by the data parallelism method, the data parallelism method can be implemented by multiple accelerators respectively executing the model training process. Among them, the multiple accelerators can be respectively stored in multiple training devices (311, 312, 313, 314), that is, Figure 8 the training device 220 in the appendix. Each training device contains one or more accelerators, and different training devices can be interconnected through the switch 300; the multiple accelerators can be stored in a certain training device to implement. Here, taking the training device 311 as an example, in the training device 311, there are a CPU 321, a memory 322, a bus 323, and n accelerators. Among them, the n accelerators include accelerator 1 (324), accelerator 2 (325),..., accelerator n (326), and n is an integer greater than 1.

[0094] In one complete training iteration of the training device 311 for training the neural network model, n accelerators respectively read different input data from the memory 322 through the bus 323. Optionally, the CPU 321 can perform a preprocessing process on the input data. Taking NLP processing as an example, the input data can include text data. The text data obtained at one time contains multiple sentences. After the CPU 321 reads it, data preprocessing is performed. Since data parallelism is required, the CPU 321 will control the input data sent to each accelerator to be different; thereafter, the specific implementation process between the CPU 321 and the n accelerators in the training device 311 can refer to Figure 1 the interaction process between the CPU 100 and the n accelerators (101, 102, 103, 104) in

[0095] Specifically, as Figure 1 shown, this training process is implemented through the CPU 100 and multiple accelerators (Accelerator 1, Accelerator 2,..., Accelerator n). The multiple accelerators jointly train. Generally, the training process includes the following steps:

[0096] 1) Create the same training model on each accelerator. For example, when training the Bert network, a complete Bert model is required on each accelerator;

[0097] 2) Initialize the weight coefficients on Accelerator 1 and send the weight coefficients to each accelerator through the broadcast operation in collective communication (1001). Generally, when training a neural network from scratch, a random method can be used to assign an initial value to the weights. In order to keep the initial weight values on each accelerator consistent, the method of randomly initializing the weights on any one accelerator first and then sending the weights to each accelerator is adopted;

[0098] 3) The CPU sends different data to different accelerators;

[0099] 4) Each accelerator performs forward and backward calculations to obtain the corresponding gradient information. This step is an operation inside each accelerator. After forward and backward calculations, the gradient information corresponding to the batch data of the current accelerator is obtained. Since the input data is ensured to be different in 3), the gradients obtained by each accelerator are different;

[0100] 5) Perform an allreduce operation to obtain the average gradient. Average the different gradients obtained by each accelerator in 4). After this step, the gradient values on all accelerators will be consistent, which are the values after averaging the gradients of each accelerator;

[0101] 6) Update the initial weights using the average gradient value. Since the operation in 2) ensures that the initial weights on each accelerator are consistent, and the update amount on each accelerator is the average gradient after being averaged by allreduce. Therefore, it can be ensured that the weight values on each accelerator can also be kept consistent after each update. Among them, in the update process of 6), each accelerator can further use the average gradient value and the weights obtained in 1) as the input of the optimizer, and perform optimization operations through the initial variables (1002) of the optimizer. The optimizer outputs the processed gradient, and each accelerator further uses the processed gradient to update the weights.

[0102] Among them, the above accelerators can be a graphics processing unit (GPU), a neural processing unit (NPU), or a tensor processing unit (TPU); gradient aggregation can be implemented in various ways, such as collective communication.

[0103] In the above data parallel processing method, training parameters such as the initial weights in step 2) and the initial variables in step 6) consume the storage space used in accelerator training. For example, the video memory used in GPU training and the memory used in CPU training. When training a larger neural network model, there will be a problem that the storage space in a single accelerator is insufficient, resulting in the inability to perform training. In this data parallel processing method, there are the following deficiencies:

[0104] 1: In step 1001, the video memory occupied by the weight coefficients of the neural network model is relatively large. In the training device of the neural network model, each of the n accelerators needs to store a copy of the weight coefficients. For small and medium-sized models, this video memory consumption is not too much. However, for large models such as Bert and GPT-2, the weight coefficients themselves occupy a large amount of video memory, and when the weight coefficients are updated, the same update operation needs to be executed on each accelerator;

[0105] 2: In step 1002, the initial variables of the optimizer occupy a large amount of video memory. In the training device of the neural network model, there are n copies of the initial variables that need to be stored in the n accelerators. The initial variables always exist in the training process and participate in iterative operations, always occupying the video memory of each accelerator. Moreover, such initial variables exist on each of the n accelerators and have equal values.

[0106] Exemplarily, taking the most typical one - layer fully - connected layer (matrix multiplication) as an example, where the number of accelerators in the training device is 4, the input layer in the table can be the output (feature map) of the previous layer, and the size and video memory consumption of this layer are shown in Table 1.

[0107]

[0108] Table 1

[0109] As can be seen from the above, in this data - parallel processing method, the weight coefficients of the neural network model and the initial variables of the optimizer are repeatedly stored on each of the n accelerators. This results in unnecessary video memory waste in the n accelerators during the model training process, that is, the video memory space utilization rate of the n accelerators is relatively low. Therefore, when training a relatively large - scale neural network model, it is often prone to the problem that the video memory space in the accelerator is insufficient, resulting in the inability to perform training.

[0110] Due to the fact that in the above - mentioned problem, both the weight coefficients of the neural network model and the initial variables of the optimizer occupy too much video memory. In the embodiments of the present application, it is considered to implement these two parameters through distributed storage in multiple accelerators of the training device. Specifically, refer to Figure 10 with respect to Figure 1 the implementation process, there are two aspects of improvement. On the one hand, step 1003 is used to replace step 1001 to perform distributed storage of the complete weight coefficients of the neural network model, and the complete weight coefficients are obtained through the collective communication method among multiple accelerators in step 1003. On the other hand, step 1004 is used to replace step 1002 to perform distributed storage of the complete initial variables of the optimizer, and the complete initial variables are obtained through the collective communication method among multiple accelerators in step 1004. Below, a training method for a neural network model in the embodiments of the present application will be introduced through the improvements in these two aspects respectively.

[0111] I. Distributed storage of the complete weight coefficients of the neural network model

[0112] Please refer to Figure 11 An embodiment of a training method for a neural network model in the embodiments of the present application includes:

[0113] 1101. Store partial weight coefficients;

[0114] In this embodiment, the training device includes multiple accelerators, and each accelerator stores partial weight coefficients in step 1101. Among them, the partial weight coefficients stored by each of the multiple accelerators in the training device together constitute the complete weight coefficients for neural network model training.

[0115] Specifically, during the training process of the neural network model, an operation for initializing the weight coefficients of the neural network model needs to be performed in the first training iteration. Among them, this initialization operation can be specified to be executed in the first accelerator, and the first accelerator is any one of multiple accelerators in the training device. After the operation of initializing the weight coefficients, the first accelerator can obtain the complete weight coefficients. Thereafter, the first accelerator divides the complete weight coefficients and sends them to multiple accelerators in the training device, realizing the distributed storage of the complete weight coefficients. That is to say, only the initialization operation of the weight coefficients needs to be performed on any one of the multiple accelerators, which can save the computing consumption of other accelerators in the training device.

[0116] In a specific implementation manner, the first accelerator can also evenly divide the complete weight coefficients and send them to multiple accelerators in the training device. Thus, in step 1101, the partial weight coefficients stored in each accelerator include the weight coefficients obtained by evenly dividing the complete weight coefficients and distributing them one by one to the multiple accelerators. During the training process of the neural network model, the processing capabilities of multiple accelerators in the training device are generally the same or nearly the same. Therefore, the complete weight coefficients can be evenly divided according to the number of multiple accelerators, and then distributed one by one to the multiple accelerators, so that each accelerator stores the evenly divided partial weight coefficients in a distributed manner. During the subsequent model training process, each accelerator uses the evenly divided weight coefficients to participate in the model training, so that the processing progress of different accelerators in the training device remains synchronized. In addition, the number of accelerators in the multiple accelerators can specifically be 2, 4, 8, 16, 32, etc., which is not limited here.

[0117] 1102. Aggregate the partial weight coefficients respectively stored in the multiple accelerators to obtain the complete weight coefficients;

[0118] In this embodiment, each accelerator in the training device aggregates the partial weight coefficients respectively stored in the multiple accelerators to obtain the complete weight coefficients.

[0119] In a specific implementation manner, in step 1102, when each accelerator in the training device aggregates the partial weight coefficients respectively stored in the multiple accelerators to obtain the complete weight coefficients, specifically, each accelerator can aggregate the partial weight coefficients respectively stored in the multiple accelerators through the gather (Allgather) operation in the collective communication mode to obtain the complete weight coefficients. The specific communication process of using Allgether in each accelerator can refer to the foregoing Figure 6The process described above will not be elaborated here. Each accelerator in the training device can obtain the complete weight coefficient through the Allgather operation in the collective communication method among multiple accelerators. This implementation provides a specific implementation process for obtaining the complete weight coefficient, improving the feasibility of the solution, and thus enhancing the implementation flexibility of this solution. In addition, during the execution of step 1102, in addition to using the collective communication method, the complete weight coefficient can also be obtained through the interaction between multiple accelerators and the CPU, or through other communication methods among multiple accelerators, which is not limited here.

[0120] 1103. Train a neural network model according to the input data and the complete weight coefficient;

[0121] In this embodiment, each accelerator in the training device trains a neural network model according to the input data and the complete weight coefficient obtained in step 1102, where the input data of multiple accelerators are different from each other.

[0122] Specifically, in the data parallel system architecture formed by multiple accelerators in the training device, a neural network training model can be created in advance in each accelerator. During one iteration training process, after obtaining the complete weight coefficient in step 1102, use the complete weight coefficient and different input data as the input of the neural network model to train the neural network model. The specific training process can refer to the Figure 2 and Figure 3 implementation process described above. After that, the output of the neural network model is gradient information, and then this gradient information can be used to update the weight coefficients stored on each accelerator (which can be the partial weight coefficients in step 1101 or the complete weight coefficient obtained in step 1102), completing one iteration training process. The multiple accelerators in this training device can execute steps 1101 to 1103 multiple times to achieve multiple iterations of the neural network model and finally obtain a better solution.

[0123] In this embodiment, during the parallel processing of training a neural network model using multiple accelerators in the training device, the complete weight coefficient in the neural network model is distributedly stored in multiple accelerators in the training device. Subsequently, the complete weight coefficient is obtained through aggregation among multiple accelerators, and then on each accelerator, the neural network model is further trained according to different input data and this complete weight coefficient. That is, by distributedly storing the complete weight coefficient in multiple accelerators in the training device, the video memory consumption of the training device during the neural network model training process is reduced.

[0124] In Figure 11In the corresponding embodiment, the complete weight coefficients of the distributed storage neural network model are specifically used to reduce a part of the video memory consumption in the accelerator. On this basis, if there is an optimizer participating in the optimization process of gradient information, the complete initial variables of the optimizer can be further distributedly stored, so as to further reduce the video memory consumption in the accelerator. The following will be based on Figure 11 the embodiment to further optimize the process of training the neural network model according to the input data and the complete weight coefficients in step 1103, and will be described through Figure 12 the specific embodiments in

[0125] Please refer to Figure 12 , another embodiment of a method for training a neural network model in an embodiment of the present application includes:

[0126] 1201. Store partial weight coefficients;

[0127] In this embodiment, the training device includes multiple accelerators, and each accelerator stores partial weight coefficients in step 1201, wherein the partial weight coefficients stored by the multiple accelerators in the training device together form the complete weight coefficients for training the neural network model.

[0128] 1202. Aggregate the partial weight coefficients respectively stored in the multiple accelerators to obtain the complete weight coefficients;

[0129] In this embodiment, each accelerator in the training device aggregates the partial weight coefficients respectively stored in the multiple accelerators to obtain the complete weight coefficients.

[0130] In this embodiment, the implementation processes of steps 1201 and 1202 can refer to the implementation processes of steps 1101 and 1102 in the foregoing Figure 11 and will not be elaborated here.

[0131] 1203. Calculate gradient information according to the input data and the complete weight coefficients;

[0132] In this embodiment, each accelerator in the training device calculates gradient information according to the input data and the complete weight coefficients obtained in step 1202, wherein the input data of the multiple accelerators are different from each other.

[0133] Specifically, in the data parallel system architecture formed by the multiple accelerators in the training device, a neural network training model can be pre-created in each accelerator. In one iteration training process, after obtaining the complete weight coefficients in step 1202, use the complete weight coefficients and different input data as the input of the neural network model to train the neural network model. The specific training process can refer to the foregoing Figure 2 and Figure 3The implementation process, after which the output of the neural network model is the gradient information.

[0134] 1204. Calculate the target gradient according to the gradient information of the multiple accelerators;

[0135] In this embodiment, each accelerator in the training device calculates the target gradient according to the gradient information obtained in step 1203, where the target gradient is used to update the partial weight coefficients stored in each accelerator in step 1201 to complete the iteration process.

[0136] In a specific implementation manner, in step 1204, when each accelerator in the training device calculates the target gradient according to the gradient information of the multiple accelerators, specifically, the target gradient can be calculated by a ReduceScatter operation in the collective communication method according to the gradient information of the multiple accelerators. The implementation process of ReduceScatter can refer to the Figure 7 implementation process described above and will not be elaborated here. Each accelerator in the training device can calculate the target gradient through a ReduceScatter operation in the collective communication method among the multiple accelerators. This implementation manner describes the specific implementation process of calculating the target gradient, improves the feasibility of the solution, and thus enhances the implementation flexibility of the present solution. In addition, during the execution of step 1204, in addition to being implemented using the collective communication method, the target gradient can also be obtained through the interaction between the multiple accelerators and the CPU, or the target gradient can be obtained through other communication methods among the multiple accelerators, which is not limited here.

[0137] 1205. Store some initial variables in the optimizer;

[0138] In this embodiment, the training device includes multiple accelerators, and each accelerator stores some initial variables of the optimizer in step 1205, where the partial initial variables stored by the multiple accelerators in the training device together form the complete initial variables of the optimizer.

[0139] In a specific implementation, in step 1205 of each accelerator in the training device, the partial initial variables stored in each of the multiple accelerators form the complete initial variables of the optimizer, where the optimizer is used to update the weight coefficients of the neural network model. Specifically, during the training process of the neural network model, an operation for initializing the initial variables of the optimizer needs to be performed during the first training iteration. Among them, the initialization operation can be specified to be executed in the first accelerator, and the first accelerator is any one of the multiple accelerators in the training device. After the operation for initializing the initial variables, the first accelerator can obtain the complete initial variables of the optimizer. Thereafter, the first accelerator splits the complete initial variables of the optimizer and sends them to the multiple accelerators in the training device, realizing the distributed storage of the complete weight coefficients. That is to say, only the initialization operation of the initial variables of the optimizer needs to be performed in any one of the multiple accelerators, which can save the computing consumption of other accelerators in the training device.

[0140] In addition, the first accelerator can also evenly split the complete initial variables of the optimizer and send them to the multiple accelerators in the training device. Thus, in step 1205, the partial weight coefficients stored in each accelerator include the initial variables obtained by evenly dividing the complete initial variables and distributing them one by one to the multiple accelerators. Among them, during the training process of the neural network model, the processing capabilities of the multiple accelerators in the training device are generally the same or nearly the same. Therefore, the complete initial variables can be evenly divided according to the number of multiple accelerators and distributed one by one to the multiple accelerators, so that each accelerator stores the evenly divided partial initial variables in a distributed manner. During the subsequent model training process, each accelerator uses the evenly divided initial variables to participate in the model training, so that the processing progress of different accelerators in the training device remains synchronized.

[0141] 1206. Process the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient;

[0142] In this embodiment, each accelerator in the training device processes the target gradient calculated in step 1204 and the partial weight coefficients stored in step 1201 according to the partial initial variables stored in step 1205 to obtain the processed target gradient.

[0143] Specifically, after the target gradient is calculated in step 1204, the target gradient can be optimized and adjusted. There are many adjustment strategies, that is, optimizers. For example, the basic one is stochastic gradient descent, which is to simply multiply the obtained target gradient by a preset learning rate and then adjust the weight coefficients stored in each accelerator. Currently, there are two widely used optimizers. One is the element-wise operation on the initial variables. After the initial variables are split into each accelerator, the calculation can be completed without communication between accelerators. For example, Adam optimizer, momentum optimizer, RMSprop optimizer, AdaMax optimizer, Adagrad optimizer, etc.; the other is the vector operation on the initial variables, including matrix or vector operations, etc. Such optimizers need to insert additional operation processes to complete the calculation during the calculation of the initial variables. For example, Lars optimizer.

[0144] For optimizers of the element-wise operation type, the operations in such optimizers are all bitwise operations and do not require operations such as partial sum and accumulation of matrices. The operations of each accelerator in step 1206 can directly perform optimization operations on the target gradient obtained locally;

[0145] For optimizers with vector operation operations, the operations in such optimizers need to perform vector operations such as matrices or vectors. Such optimizers require the participation of the complete gradient during the calculation. Since the target gradient is distributed on each accelerator, additional collective communication needs to be added during the operations of each accelerator in step 1206 to ensure the correctness of the calculation. The implementation process of this type of optimizer will be described in detail below:

[0146] Specifically, in step 1206, if the optimizer includes vector operations, when each accelerator in the training device processes the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient, it can specifically calculate the scalar representation of the target gradient; then aggregate the scalar representations of the target gradient in the multiple accelerators to obtain the summation result of the target gradient; thereafter, each accelerator calculates the vector representation of the target gradient according to the summation result; and further processes the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient. Among them, if the optimizer includes appropriate operations (such as matrix operations, vector operations, or other vector operations, etc.), complete gradients are required to participate in the calculation. Since the target gradient is distributed on each accelerator, when each accelerator uses the accelerator to obtain the processed target gradient, each accelerator first calculates the scalar representation of the target gradient; then aggregates the scalar representations of the target gradient in the multiple accelerators to obtain the summation result of the target gradient; thereafter, each accelerator calculates the vector representation of the target gradient according to the summation result; and further processes the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient. In this implementation, the solution can be applied to the implementation process of an optimizer containing vector operations, thereby improving the feasibility of the solution.

[0147] Exemplarily, taking Figure 13 the data structure in the shown Lars optimizer as an example, when calculating the norm parameter in the Lars optimizer, the target gradients obtained by each accelerator in step 1301 are {X 0 , X 1 , X 2 ...X n}, where Xi is a one-dimensional or multi-dimensional tensor, and i is any integer from 1 to n; the process executed by each accelerator in step 1206 can be Figure 13 the process of squaring Xi in 0 2 to obtain {X 1 2 , X 2 2 ...X n 2}; obviously, in addition to the squaring process, there can be other ways to obtain the vector representation of Xi, and there can also be other ways, such as calculating the modulus of a vector, calculating the square root value of a vector, etc., which can all obtain the vector representation, and are not limited here; thereafter, each accelerator calculates the sum of squares in step 1303 to obtain P_S = X 0 2 + X 12 +X 2 2 +...+X n 2 ; Further, in the process of collective communication of each accelerator in step 1304, specifically, the vector representation of the target gradient can be processed according to the parameter P_S and other first initial variables in the Lars optimizer to obtain the processed target gradient.

[0148] In addition, in the above operation, when each accelerator in the training device aggregates the scalar representations of the target gradient in the multiple accelerators to obtain the sum result of the target gradient, specifically, the scalar representations of the target gradient in the multiple accelerators can be aggregated through the reduce operation (Allreduce) in the collective communication method to obtain the sum result of the target gradient. Among them, the specific implementation process of each accelerator using the Allreduce communication method can specifically refer to the Figure 4 implementation process described above, which will not be elaborated here. Each accelerator in the training device can obtain the sum result of the target gradient through the reduce operation (allreduce) in the collective communication method among the multiple accelerators. This implementation method provides a specific implementation process for obtaining the sum result of the target gradient, improves the feasibility of the solution, and thus improves the implementation flexibility of this solution. In addition, during the execution process of each accelerator aggregating to obtain the sum result of the target gradient, in addition to using the collective communication method, the sum result of the target gradient can also be obtained through the interaction between the multiple accelerators and the CPU, and the sum result of the target gradient can also be obtained through other communication methods among the multiple accelerators, which is not limited here.

[0149] 1207. Update the partial weight coefficients using the processed target gradient, and train the neural network model according to the updated partial weight coefficients.

[0150] In this embodiment, each accelerator in the training device updates the partial weight coefficients stored in step 1201 using the processed target gradient obtained in step 1206, and trains the neural network model according to the updated partial weight coefficients.

[0151] Specifically, in step 1207, when each accelerator in the training device updates the partial weight coefficients using the target gradient, it can specifically update the partial weight coefficients stored in step 1201 according to the processed target gradient obtained in step 1206. When using an optimizer to optimize the weight coefficients of the neural network model, each accelerator can perform distributed storage of the initial variables of the accelerator, that is, store partial initial variables in each accelerator. Then each accelerator uses the target gradient and the partial weight coefficients as the input of the optimizer, and after being optimized by the preset optimization algorithm in the optimizer, the processed target gradient is obtained. Then, according to the processed target gradient, the partial weight coefficients stored in each accelerator are updated, that is, the complete initial variables in the optimizer are distributed and stored in multiple accelerators in the training device, thereby further reducing the video memory consumption of the training device during the training process of the neural network model.

[0152] In addition, in a specific implementation manner, each accelerator in the training device can further obtain the update parameters of the partial weight coefficients and update the partial weight coefficients according to the update parameters of the partial weight coefficients; and / or, each accelerator is further used to obtain the update parameters of the initial variables and update the initial variables according to the update parameters of the variables; and / or, each accelerator is further used to obtain the update parameters of the target gradient and update the target gradient according to the update parameters of the target gradient; and / or, each accelerator is further used to obtain the update parameters of the processed target gradient and update the target gradient according to the update parameters of the processed target gradient. Among them, the target parameters involved in the training process of the neural network model can be distributed and stored. The target parameters include partial weight coefficients and / or initial variables and / or target gradients and / or processed target gradients, etc. Thus, when there is an update to the partial weight coefficients corresponding to the neural network model or the initial variables in the optimizer, the distributed target parameters can be updated respectively in each accelerator in the training device. As can be seen from the above, since training parameters such as complete weight coefficients and complete initial variables are distributed and stored, in the step of neural network optimization and update, the calculation can be performed only on the locally stored training parameters, thereby reducing the repeated calculation in the existing data parallel scheme, reducing the overall calculation amount, and further reducing the video memory consumption of the accelerators in the training device.

[0153] Taking the training of Bert and GPT-2 networks as an example, the number of parameters of the Bert network is about 1.3G. The training network uses the Adam optimizer. If the existing data parallel method is used for training, only the occupied space of the weight parameters and the initial variables in the optimizer is concerned, which is about 7.8GBytes in size. If it is GPT-2, it is 36GBytes in size. And most of the GPU cards on the market currently have a video memory size of 16GBytes. Coupled with the size of the feature map, it is very easy to run out of video memory and unable to train, and GPT-2 simply cannot fit. After Figure 11 and Figure 12 In the embodiments shown, compared with the implementation process of the prior art solution, the video memory optimization effect on the Bert and GPT-2 networks is as shown in Table 2. In the case of a single machine with an 8-card cluster (that is, the training device contains 8 accelerators), this situation can be fully improved. Only 0.975GBytes of video memory is consumed on a single card of Bert, and only 4.5GBytes of video memory is required for GPT-2, significantly reducing the video memory consumption of the training device during the training process of the neural network model. If the number of machines is further increased, such as when there are 32 cards (that is, the training device contains 32 accelerators), the video memory consumption of the training device during the training process of the neural network model will be further reduced.

[0154]

[0155] Table 2

[0156] In this embodiment, during the parallel processing of training a neural network model using multiple accelerators in the training device, the complete weight coefficients in the neural network model are distributed and stored in multiple accelerators in the training device. Subsequently, the complete weight coefficients are obtained through aggregation in multiple accelerators. After that, when using the optimizer to optimize the weight coefficients of the neural network model, each accelerator can distribute and store the initial variables of the accelerator, that is, each accelerator stores part of the initial variables. Each accelerator then uses the target gradient and part of the weight coefficients as the input of the optimizer, and through the preset optimization algorithm in the optimizer, the processed target gradient is obtained, and then the part of the weight coefficients stored in each accelerator is updated according to the processed target gradient, that is, by distributing and storing the complete initial variables in the optimizer in multiple accelerators in the training device, thereby further reducing the video memory consumption of the training device during the training process of the neural network model.

[0157] II. Distributed storage of the complete initial variables of the optimizer

[0158] Please refer to Figure 14 , another embodiment of a method for training a neural network model in an embodiment of the present application includes:

[0159] 1401. Calculate gradient information based on the input data and the complete weight coefficients;

[0160] In this embodiment, the training device includes multiple accelerators. Each accelerator calculates gradient information based on the input data and the complete weight coefficients in step 1401. In the system architecture with data parallelism formed by the multiple accelerators in the training device, a neural network training model can be pre-created in each accelerator. The complete weight coefficients can be obtained by initializing the weight coefficients in this neural network model in each accelerator. In addition, in step 1401, the input data of each of the multiple accelerators is different.

[0161] 1402. Calculate the target gradient based on the gradient information of the multiple accelerators;

[0162] In this embodiment, each accelerator in the training device calculates the target gradient based on the gradient information calculated in step 1401. The target gradient can be used to update the complete weight coefficients stored in each accelerator in step 1401 to complete the iteration process.

[0163] In a specific implementation, in step 1402, when each accelerator in the training device calculates the target gradient based on the gradient information of the multiple accelerators, specifically, the target gradient can be calculated through the ReduceScatter operation in the collective communication method based on the gradient information of the multiple accelerators. The implementation process of ReduceScatter can refer to the Figure 7 foregoing implementation process and will not be elaborated here. Each accelerator in the training device can calculate the target gradient through the ReduceScatter operation in the collective communication method among the multiple accelerators. This implementation mode discloses the specific implementation process of calculating the target gradient, improves the feasibility of the solution, and thus enhances the implementation flexibility of this solution. In addition, during the execution of step 1402, in addition to implementing it using the collective communication method, the target gradient can also be obtained through the interaction between the multiple accelerators and the CPU, or the target gradient can be obtained through other communication methods among the multiple accelerators, which is not limited here.

[0164] 1403. Store some initial variables in the optimizer;

[0165] In this embodiment, the training device includes multiple accelerators. Each accelerator stores some initial variables of the optimizer in step 1403. The partial initial variables stored by each of the multiple accelerators in the training device form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model.

[0166] In a specific implementation, in step 1403 of each accelerator in the training device, the partial initial variables stored in each of the multiple accelerators form the complete initial variables of the optimizer, where the optimizer is used to update the weight coefficients of the neural network model. Specifically, during the training process of the neural network model, an operation for initializing the initial variables of the optimizer needs to be performed during the first training iteration. Among them, the initialization operation can be specified to be executed in the first accelerator, and the first accelerator is any one of the multiple accelerators in the training device. After the operation for initializing the initial variables, the first accelerator can obtain the complete initial variables of the optimizer. Thereafter, the first accelerator splits the complete initial variables of the optimizer and sends them to the multiple accelerators in the training device, realizing the distributed storage of the complete weight coefficients. That is to say, only the initialization operation of the initial variables of the optimizer needs to be performed on any one of the multiple accelerators, which can save the computing consumption of other accelerators in the training device.

[0167] In addition, the first accelerator can also evenly split the complete initial variables of the optimizer and send them to the multiple accelerators in the training device. Thus, in step 1403, the partial weight coefficients stored in each accelerator include the initial variables obtained by evenly dividing the complete initial variables and distributing them one by one to the multiple accelerators. During the training process of the neural network model, the processing capabilities of the multiple accelerators in the training device are generally the same or nearly the same. Therefore, the complete initial variables can be evenly divided according to the number of multiple accelerators and distributed one by one to the multiple accelerators, so that each accelerator stores the evenly divided partial initial variables in a distributed manner. During the subsequent model training process, each accelerator uses the evenly divided initial variables to participate in the model training, so that the processing progress of different accelerators in the training device remains synchronized.

[0168] 1404. Process the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient;

[0169] In this embodiment, each accelerator in the training device processes the target gradient calculated in step 1402 and the partial weight coefficients according to the partial initial variables stored in step 1403 to obtain the processed target gradient, where the partial weight coefficients processed by each of the multiple accelerators form the complete weight coefficients in step 1401.

[0170] Specifically, after the target gradient is calculated in step 1204, the target gradient can be optimized and adjusted. There are many adjustment strategies, that is, optimizers. Basic ones such as stochastic gradient descent simply multiply the obtained target gradient by a preset learning rate and then adjust the weight coefficients stored in each accelerator. Currently, there are two widely used optimizers. One is an element-wise operation on the initial variables. After the initial variables are split into each accelerator, calculations can be completed without communication between accelerators. For example, Adam optimizer, momentum optimizer, RMSprop optimizer, AdaMax optimizer, Adagrad optimizer, etc.; the other is a vector operation on the initial variables (including matrix or vector operations, etc.). Such an optimizer requires inserting additional calculation processes to complete the calculation during the calculation of the initial variables. For example, Lars optimizer. The specific implementation processes of these two optimizers can refer to the implementation process of step 1206 described above and will not be elaborated here.

[0171] 1405. Update the complete weight coefficients according to the processed target gradient, and train the neural network model according to the updated complete weight coefficients.

[0172] In this embodiment, each accelerator in the training device updates the complete weight coefficients pre-stored in each accelerator in step 1401 according to the processed target gradient obtained in step 1404, obtains the updated complete weight coefficients, and trains the neural network model according to the updated complete weight coefficients. Among them, it can be seen from step 1401 that the data volume of the gradient information in each accelerator corresponds to the data volume of the complete weight coefficients. It can be seen from steps 1402 and 1403 that the data volume of the target gradient corresponds to the data volume of the partial weight coefficients. Therefore, in step 1404, each accelerator in the training device realizes a partial update of the complete weight coefficients.

[0173] Specifically, in the data parallel system architecture formed by multiple accelerators in the training device, a neural network training model can be created in advance in each accelerator. During one iteration training process, after obtaining the complete weight coefficients in step 1401, use the complete weight coefficients and different input data as the input of the neural network model to train the neural network model. The specific training process can refer to the foregoing Figure 2 and Figure 3During the implementation process, the output of the neural network model is then gradient information, and thereafter, this gradient information can be used to update the weight coefficients stored on each accelerator (which can be the complete weight coefficients in step 1401 or the partial weight coefficients obtained in step 1404), completing one iteration training process. The multiple accelerators in this training device can execute steps 1401 to 1405 multiple times to achieve multiple iterations of the neural network model and finally obtain a better solution.

[0174] In addition, each accelerator in the training device can further obtain the update parameters of the complete weight coefficients and update the complete weight coefficients according to the update parameters of the complete weight coefficients; and / or, each accelerator is also used to obtain the update parameters of the initial variables and update the initial variables according to the update parameters of the initial variables; and / or, each accelerator is also used to obtain the update parameters of the target gradient and update the target gradient according to the update parameters of the target gradient; and / or, each accelerator is also used to obtain the update parameters of the processed target gradient and update the target gradient according to the update parameters of the processed target gradient. Among them, the target parameters involved in the neural network model training process can be stored distributively, and the target parameters include complete weight coefficients and / or initial variables and / or target gradients and / or processed target gradients, etc. Thus, when there is an update to the partial weight coefficients corresponding to the neural network model or the initial variables in the optimizer, the distributively stored target parameters can be updated respectively in each accelerator in the training device, thereby further reducing the video memory consumption of the accelerators in the training device.

[0175] In this embodiment, during the parallel processing of training a neural network model using multiple accelerators in the training device, the complete initial variables of the optimizer in the neural network model are stored distributively in the multiple accelerators in the training device. Each accelerator then processes the target gradient and partial weight coefficients according to this partial initial variable to obtain the processed target gradient. Thereafter, each accelerator further updates the complete weight coefficients according to the processed target gradient and trains the neural network model according to the updated complete weight coefficients, that is, by distributively storing the complete initial weights of the optimizer in the multiple accelerators in the training device, thereby reducing the video memory consumption of the training device during the neural network model training process.

[0176] In Figures 1 to 14 Based on the corresponding embodiments, in order to better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions.

[0177] Specifically, please refer to Figure 15 , Figure 15A schematic structural diagram of the training device 1500 provided by an embodiment of the present application. Among them, the training device 1500 includes multiple accelerators, and specifically includes in each accelerator:

[0178] A storage unit 1501 for storing partial weight coefficients, and the partial weight coefficients stored in each of the multiple accelerators form a complete weight coefficient;

[0179] An aggregation unit 1502 for aggregating the partial weight coefficients separately stored in the multiple accelerators to obtain the complete weight coefficient;

[0180] A training unit 1503 for training a neural network model according to input data and the complete weight coefficient, wherein the input data of the multiple accelerators are different from each other.

[0181] In a possible design, the training unit 1503 is specifically used for:

[0182] Calculating gradient information according to the input data and the complete weight coefficient;

[0183] Calculating a target gradient according to the gradient information of the multiple accelerators;

[0184] Updating the partial weight coefficients by using the target gradient, and training the neural network model according to the updated partial weight coefficients.

[0185] In a possible design, the storage unit 1501 is further used for:

[0186] Storing partial initial variables in an optimizer, and the partial initial variables stored in each of the multiple accelerators form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model;

[0187] The training unit 1503 is specifically used for:

[0188] Processing the target gradient and the partial weight coefficients according to the partial initial variables to obtain a processed target gradient;

[0189] Updating the partial weight coefficients according to the processed target gradient.

[0190] In a possible design, the optimizer includes vector operations, and the training unit 1503 is specifically used for:

[0191] Calculating a scalar representation of the target gradient;

[0192] Aggregating the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient;

[0193] Calculate the vector representation of the target gradient according to the summation result;

[0194] Process the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient.

[0195] In a possible design, the training unit 1503 is specifically configured to:

[0196] Aggregate the scalar representations of the target gradients in the multiple accelerators through a reduction operation in the collective communication mode to obtain the summation result of the target gradients.

[0197] In a possible design, the partial weight coefficients include the weight coefficients obtained by evenly dividing the complete weight coefficients and distributing them to the multiple accelerators one by one.

[0198] In a possible design, the aggregation unit 1502 is specifically configured to:

[0199] Aggregate the partial weight coefficients respectively stored in the multiple accelerators through a gather operation in the collective communication mode to obtain the complete weight coefficient.

[0200] In a possible design, the training unit 1503 is specifically configured to:

[0201] Calculate the target gradient according to the gradient information of the multiple accelerators through a reduce-scatter operation in the collective communication mode.

[0202] In a possible design, the partial initial variables include the initial variables obtained by evenly dividing the complete initial variables and distributing them to the multiple accelerators one by one.

[0203] It should be noted that the information interaction, execution process, etc. among the modules / units in the training device 1500 are based on the same concept as the foregoing Figure 11 and Figure 12 The embodiments shown. For specific content, reference can be made to the descriptions in the foregoing embodiments shown in this application, which will not be elaborated here.

[0204] Specifically, please refer to Figure 16 , Figure 16 which is another structural schematic diagram of the training device 1600 provided by the embodiments of this application. Among them, the training device 1600 includes multiple accelerators, and each accelerator specifically includes:

[0205] A calculation unit 1601, configured to calculate gradient information according to input data and complete weight coefficients, where the input data of the multiple accelerators are different from each other;

[0206] The computing unit 1601 is further configured to calculate a target gradient according to the gradient information of the multiple accelerators;

[0207] The storage unit 1602 is configured to store some initial variables in the optimizer. The some initial variables stored in each of the multiple accelerators form the complete initial variables of the optimizer. The optimizer is used to update the weight coefficients of the neural network model;

[0208] The processing unit 1603 is configured to process the target gradient and some weight coefficients according to the some initial variables to obtain a processed target gradient. The some weight coefficients processed by each of the multiple accelerators form the complete weight coefficients;

[0209] The updating unit 1604 is configured to update the complete weight coefficients according to the processed target gradient, and train the neural network model according to the updated complete weight coefficients.

[0210] In a possible design, the optimizer includes vector operations. Specifically, the processing unit 1603 is configured to:

[0211] Calculate a scalar representation of the target gradient;

[0212] Aggregate the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient;

[0213] Calculate a vector representation of the target gradient according to the summation result;

[0214] Process the vector representation of the target gradient and the some weight coefficients according to the some initial variables to obtain the processed target gradient.

[0215] In a possible design, specifically, the processing unit 1603 is configured to:

[0216] Aggregate the scalar representations of the target gradient in the multiple accelerators through a reduction operation in the collective communication mode to obtain a summation result of the target gradient.

[0217] In a possible design, specifically, the computing unit 1601 is configured to:

[0218] Calculate the target gradient according to the gradient information of the multiple accelerators through a reduce-scatter operation in the collective communication mode.

[0219] In a possible design, the some initial variables include the initial variables obtained by evenly dividing the complete initial variables and allocating them to the multiple accelerators one by one.

[0220] It should be noted that the information interaction, execution process, etc. between the modules / units in the training device 1600 are based on the same concept as the corresponding embodiments in this application. For specific content, reference can be made to the descriptions in the foregoing embodiments shown in this application, and details will not be repeated here. Figure 14 The corresponding embodiments are based on the same concept, and for specific content, reference can be made to the descriptions in the foregoing embodiments shown in this application, and details will not be repeated here.

[0221] An embodiment of this application also provides a training device. Please refer to Figure 17 , Figure 17 which is a schematic structural diagram of a training device provided by an embodiment of this application. The training device 1700 can be deployed with Figure 15 the training device 1500 of the neural network described in the corresponding embodiment, for implementing Figure 11 and Figure 12 the functions of the training device in the corresponding embodiment. Alternatively, the training device 1700 can be deployed with Figure 16 the training device 1600 of the neural network described in the corresponding embodiment, for implementing Figure 14 the functions of the training device in the corresponding embodiment. Specifically, the training device 1700 is implemented by one or more training devices. The training device 1700 may have relatively large differences due to different configurations or performances. It may include one or more central processing units (CPUs) 1722 (for example, one or more processors) and a memory 1732, and one or more storage media 1730 (for example, one or more mass storage devices) for storing application programs 1742 or data 1744. Among them, the memory 1732 and the storage media 1730 can be transient storage or persistent storage. The program stored in the storage media 1730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations for the training device. Further, the central processor 1722 can be set to communicate with the storage media 1730 and execute a series of instruction operations in the storage media 1730 on the training device 1700. However, it should be understood that Figure 17 the training device shown in

[0222] The training device 1700 may also include one or more power supplies 1726, one or more wired or wireless network interfaces 1750, one or more input / output interfaces 1758, and / or one or more operating systems 1741, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0223] In an embodiment of the present application, the central processing unit 1722 is configured to execute Figure 3 the method for obtaining a neural network executed by the training device in the corresponding embodiment, or is configured to execute Figure 17 the method for obtaining a neural network executed by the training device in the corresponding embodiment. It should be noted that for the specific implementation manner of the central processing unit 1722 executing the method for obtaining a neural network, reference can be made to Figure 3 and Figure 17 the descriptions in the corresponding method embodiments, which will not be elaborated here one by one.

[0224] In an embodiment of the present application, there is also provided a computer program product, which when running on a computer, causes the computer to execute the steps executed by the training device in the method described in the foregoing Figure 11 and Figure 12 embodiments, or causes the computer to execute the steps executed by the training device in the method described in the foregoing Figure 14 embodiments.

[0225] In an embodiment of the present application, there is also provided a computer-readable storage medium, in which a program for signal processing is stored, which when running on a computer, causes the computer to execute the steps executed by the training device in the method described in the foregoing Figure 11 and Figure 12 embodiments, or causes the computer to execute the steps executed by the training device in the method described in the foregoing Figure 14 embodiments.

[0226] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0227] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CLUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of this application, in essence or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.

[0228] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0229] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)), etc.

Claims

1. A training device for a neural network model, characterized in that, the training device includes a plurality of accelerators, and each accelerator is configured to: store partial weight coefficients, and the partial weight coefficients stored by each of the plurality of accelerators together form a complete weight coefficient; aggregate the partial weight coefficients respectively stored in the plurality of accelerators to obtain the complete weight coefficient; train the neural network model according to input data and the complete weight coefficient, wherein the input data of the plurality of accelerators are different from each other, and the input data is text data.

2. The training device according to claim 1, characterized in that, when each accelerator trains the neural network model according to the input data and the complete weight coefficient, it is specifically configured to: calculate gradient information according to the input data and the complete weight coefficient; calculate a target gradient according to the gradient information of the plurality of accelerators; update the partial weight coefficients by using the target gradient, and train the neural network model according to the updated partial weight coefficients.

3. The training device according to claim 2, characterized in that, each accelerator is configured to: store partial initial variables in an optimizer, and the partial initial variables stored by each of the plurality of accelerators together form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model; when each accelerator updates the partial weight coefficients by using the target gradient, it is specifically configured to: process the target gradient and the partial weight coefficients according to the partial initial variables to obtain a processed target gradient; update the partial weight coefficients according to the processed target gradient.

4. The training device according to claim 3, characterized in that, the optimizer includes vector operations, and when each accelerator processes the target gradient and the partial weight coefficients according to the partial initial variables to obtain a processed target gradient, it is specifically configured to: calculate a scalar representation of the target gradient; aggregate the scalar representations of the target gradients in the plurality of accelerators to obtain a summation result of the target gradients; calculate a vector representation of the target gradient according to the summation result; process the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient.

5. The training device according to claim 4, characterized in that, when each accelerator aggregates the scalar representations of the target gradients in the plurality of accelerators to obtain a summation result of the target gradients, it is specifically configured to: aggregate the scalar representations of the target gradients in the plurality of accelerators through a reduction operation in a collective communication manner to obtain a summation result of the target gradients.

6. The training device according to any one of claims 1 to 5, characterized in that, the partial weight coefficients include weight coefficients obtained by equally dividing the complete weight coefficient and distributing them to the plurality of accelerators one by one.

7. The training device according to any one of claims 1 to 5, characterized in that, When each of the accelerators aggregates the partial weight coefficients respectively stored in the multiple accelerators to obtain the complete weight coefficient, it specifically is used for: Aggregating the partial weight coefficients respectively stored in the multiple accelerators through a gather operation in the collective communication mode to obtain the complete weight coefficient.

8. The training device according to any one of claims 2 to 5, wherein, When each of the accelerators calculates the target gradient according to the gradient information of the multiple accelerators, it specifically is used for: Calculating the target gradient according to the gradient information of the multiple accelerators through a reduce scatter operation in the collective communication mode.

9. The training device according to claim 3 or 4, wherein, The partial initial variables include the initial variables obtained by evenly dividing the complete initial variables and distributing them to the multiple accelerators one by one.

10. A training device for a neural network model, wherein, The training device includes multiple accelerators, and each accelerator is used for: Calculating gradient information according to input data and the complete weight coefficient, wherein the input data of the multiple accelerators are different from each other, and the input data is text data; Calculating the target gradient according to the gradient information of the multiple accelerators; Storing partial initial variables in an optimizer, and the partial initial variables respectively stored by the multiple accelerators form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model; Processing the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient, and the partial weight coefficients respectively processed by the multiple accelerators form the complete weight coefficient; Updating the complete weight coefficient according to the processed target gradient, and training the neural network model according to the updated complete weight coefficient.

11. The training device according to claim 10, wherein, The optimizer includes vector operations. When each of the accelerators processes the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient, it specifically is used for: Calculating a scalar representation of the target gradient; Aggregating the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient; Calculating a vector representation of the target gradient according to the summation result; Processing the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient.

12. The training device according to claim 11, wherein, When each of the accelerators aggregates the scalar representations of the target gradient in the multiple accelerators to obtain the summation result of the target gradient, it specifically is used for: Aggregating the scalar representations of the target gradient in the multiple accelerators through a reduce operation in the collective communication mode to obtain the summation result of the target gradient.

13. The training device according to any one of claims 10 to 12, wherein, When each of the accelerators calculates the target gradient according to the gradient information of the multiple accelerators, it specifically is used for: Calculate the target gradient according to the gradient information of the multiple accelerators through a reduction operation in the collective communication mode.

14. The training device according to any one of claims 10 to 12, wherein, The partial initial variables include the initial variables obtained by evenly dividing the complete initial variables and distributing them to the multiple accelerators one by one.

15. A training method for a neural network model, wherein, The training method is applied to multiple accelerators, and the multiple accelerators are included in a training device. The method includes: Store partial weight coefficients, and the partial weight coefficients stored by each of the multiple accelerators form a complete weight coefficient; Aggregate the partial weight coefficients stored separately in the multiple accelerators to obtain the complete weight coefficient; Train a neural network model according to input data and the complete weight coefficient, wherein the input data of the multiple accelerators are different from each other, and the input data is text data.

16. The training method according to claim 15, wherein, The training of the neural network model according to the input data and the complete weight coefficient includes: Calculate gradient information according to the input data and the complete weight coefficient; Calculate a target gradient according to the gradient information of the multiple accelerators; Update the partial weight coefficients using the target gradient, and train the neural network model according to the updated partial weight coefficients.

17. The training method according to claim 16, wherein, The method further includes: Store partial initial variables in an optimizer, and the partial initial variables stored by each of the multiple accelerators form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model; The updating of the partial weight coefficients using the target gradient includes: Process the target gradient and the partial weight coefficients according to the partial initial variables to obtain a processed target gradient; Update the partial weight coefficients according to the processed target gradient.

18. The training method according to claim 17, wherein, The optimizer includes vector operations. The processing of the target gradient and the partial weight coefficients according to the partial initial variables to obtain a processed target gradient includes: Calculate a scalar representation of the target gradient; Aggregate the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient; Calculate a vector representation of the target gradient according to the summation result; Process the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient.

19. The training method according to claim 18, wherein, The aggregating of the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient includes: Aggregate the scalar representations of the target gradient in the multiple accelerators through a reduction operation in the collective communication mode to obtain a summation result of the target gradient.

20. The training method according to any one of claims 15 to 19, wherein, The partial weight coefficients include the weight coefficients obtained by evenly dividing the complete weight coefficients and distributing them one by one to the multiple accelerators.

21. The training method according to any one of claims 15 to 19, wherein, the step of aggregating the partial weight coefficients respectively stored in the multiple accelerators to obtain the complete weight coefficients includes: aggregating the partial weight coefficients respectively stored in the multiple accelerators through a gather operation in the collective communication method to obtain the complete weight coefficients.

22. The training method according to any one of claims 16 to 19, wherein, the step of calculating the target gradient according to the gradient information of the multiple accelerators includes: calculating the target gradient according to the gradient information of the multiple accelerators through a reduce-scatter operation in the collective communication method.

23. The training method according to claim 17 or 18, wherein, the partial initial variables include the initial variables obtained by evenly dividing the complete initial variables and distributing them one by one to the multiple accelerators.

24. A training method for a neural network model, wherein, the training method is applied to multiple accelerators, and the multiple accelerators are included in a training device. The method includes: calculating gradient information according to input data and complete weight coefficients, wherein the input data of the multiple accelerators are different from each other, and the input data is text data; calculating a target gradient according to the gradient information of the multiple accelerators; storing partial initial variables in an optimizer, and the partial initial variables respectively stored in the multiple accelerators form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model; processing the target gradient and partial weight coefficients according to the partial initial variables to obtain a processed target gradient, and the partial weight coefficients respectively processed by the multiple accelerators form the complete weight coefficients; updating the complete weight coefficients according to the processed target gradient, and training the neural network model according to the updated complete weight coefficients.

25. The training method according to claim 24, wherein, the optimizer includes vector operations, and the step of processing the target gradient and partial weight coefficients according to the partial initial variables to obtain a processed target gradient includes: calculating a scalar representation of the target gradient; aggregating the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient; calculating a vector representation of the target gradient according to the summation result; processing the vector representation of the target gradient and the partial weight coefficients according to the partial initial variables to obtain the processed target gradient.

26. The training method according to claim 25, wherein, the step of aggregating the scalar representations of the target gradient in the multiple accelerators to obtain a summation result of the target gradient includes: aggregating the scalar representations of the target gradient in the multiple accelerators through a reduce operation in the collective communication method to obtain a summation result of the target gradient.

27. The training method according to any one of claims 24 to 26, wherein, Calculating the target gradient according to the gradient information of the multiple accelerators includes: Calculating the target gradient according to the gradient information of the multiple accelerators through a reduce-scatter operation in the collective communication mode.

28. The training method according to any one of claims 24 to 26, wherein, The partial initial variables include the initial variables obtained by evenly dividing the complete initial variables and distributing them to the multiple accelerators one by one.

29. A training device for a neural network model, wherein, The training device includes multiple accelerators, and each accelerator includes: A storage unit for storing partial weight coefficients, and the partial weight coefficients stored by each of the multiple accelerators form a complete weight coefficient; An aggregation unit for aggregating the partial weight coefficients separately stored in the multiple accelerators to obtain the complete weight coefficient; A training unit for training a neural network model according to input data and the complete weight coefficient, wherein the input data of the multiple accelerators are different from each other, and the input data is text data.

30. A training device for a neural network model, wherein, The training device includes multiple accelerators, and each accelerator includes: A calculation unit for calculating gradient information according to input data and a complete weight coefficient, wherein the input data of the multiple accelerators are different from each other, and the input data is text data; The calculation unit is further configured to calculate a target gradient according to the gradient information of the multiple accelerators; A storage unit for storing partial initial variables in an optimizer, and the partial initial variables stored by each of the multiple accelerators form the complete initial variables of the optimizer, and the optimizer is used to update the weight coefficients of the neural network model; A processing unit for processing the target gradient and partial weight coefficients according to the partial initial variables to obtain a processed target gradient, and the partial weight coefficients processed by each of the multiple accelerators form the complete weight coefficient; An update unit for updating the complete weight coefficient according to the processed target gradient and training the neural network model according to the updated complete weight coefficient.

31. A computer-readable storage medium, wherein, The computer storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method according to any one of claims 15 to 23, or when the program instructions are executed by a processor, the processor executes the method according to any one of claims 24 to 28.

32. A chip, wherein, The chip includes a processor and a data interface, and the processor reads instructions stored on a memory through the data interface to execute the method according to any one of claims 15 to 23, or to execute the method according to any one of claims 24 to 28.

Citation Information

Patent Citations

  • Neural network training and image processing method and device, equipment and medium

    CN110363297A