Method and device for training large language model
By normalizing the target parameter matrix of the output layer of the large language model and mapping update, the problem of training instability is solved and the training stability of the large language model is improved.
Patent Information
- Application Number
- CN202510225377.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
When training large language models, due to the large amount of model parameters, training instability often occurs, such as the sudden increase in training losses, resulting in a decrease in training stability.
By normalizing the target parameter matrix of the output layer, the normalized target parameter matrix is obtained, and the input vector is mapped using this matrix to update the original parameter matrix and improve training stability.
Through the normalization of the output layer, the numerical range of the output results of the output layer can be effectively controlled, the training stability of the large language model can be improved, and the problem of sudden increase in training losses can be avoided.
Smart Images

Figure CN120068972A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of artificial intelligence, and in particular, to a method and apparatus for training large language models. Background Art
[0002] In the field of artificial intelligence, large language models refer to models with a large number of parameters. For example, deep neural networks with more than 10B parameters can process massive amounts of data and complete various complex tasks, such as natural language processing, computer vision, speech recognition, etc. With the continuous improvement of computer hardware performance and the continuous optimization of deep learning algorithms, the development of large language models has become increasingly rapid. The parameter scale of large language models has been continuously expanding, the training time has been getting longer, and the performance has also been improved. Now, large language models have become one of the important research directions in the field of artificial intelligence, and many enterprises and institutions are researching and developing their own large language models in order to achieve better performance in various tasks.
[0003] Currently, when training large language models, due to the excessive number of model parameters, problems such as unstable training are often encountered. For example, the training loss suddenly increases. Therefore, a solution is needed to improve the stability of large language model training. Summary of the Invention
[0004] One or more embodiments of this specification describe a method for training large language models, which can improve the stability of large language model training.
[0005] In a first aspect, a method for training a large language model is provided. The large language model includes an output layer. The method includes:
[0006] Obtain a target parameter matrix of the output layer, which is obtained by normalizing the original parameter matrix obtained by the output layer in the previous batch of training;
[0007] Through the target parameter matrix, perform mapping processing on the input vector of the output layer to obtain an output result mapped to a preset vocabulary space; the input vector corresponds to the input text;
[0008] After obtaining the output results corresponding to each input text in each micro-batch included in the current batch, determine the target parameter gradient, and use the target parameter gradient to update the original parameter matrix.
[0009] In a second aspect, an apparatus for training a large language model is provided. The large language model includes an output layer. The apparatus includes:
[0010] An acquisition unit for acquiring a target parameter matrix of the output layer, which is obtained by normalizing an original parameter matrix obtained by the output layer in the previous batch of training;
[0011] A mapping unit for mapping an input vector of the output layer through the target parameter matrix to obtain an output result mapped to a preset vocabulary space; the input vector corresponds to the input text;
[0012] A determination unit for determining a target parameter gradient after obtaining the output results corresponding to the respective input texts in each micro-batch included in the current batch, and updating the original parameter matrix using the target parameter gradient.
[0013] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.
[0014] In a fourth aspect, a computing device is provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method of the first aspect is implemented.
[0015] The method and apparatus for training a large language model provided by one or more embodiments of this specification map the input vector of the output layer through the parameters of the normalization process of the output layer. Thus, the numerical range of the output result of the output layer can be effectively controlled, which helps to improve the stability of the training of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 Showing a schematic diagram of a distributed training framework in an example of this specification;
[0018] Figure 2 Showing a flowchart of a method for training a large language model according to an embodiment of this specification;
[0019] Figure 3 Showing a schematic diagram of a normalization process method in an example of this specification;
[0020] Figure 4 Showing a schematic diagram of an apparatus for training a large language model according to an embodiment of this specification. DETAILED DESCRIPTION
[0021] The following describes the solution provided in this specification in conjunction with the accompanying drawings.
[0022] As mentioned above, during the process of training large language models, problems of unstable training may occur. Currently, there are the following two solutions:
[0023] The first one is Gradient Clipping, that is, when the norm of the gradient (usually the L2 norm) exceeds a preset threshold, the gradient is scaled so that its norm does not exceed this threshold. However, this method is a passive correction method and cannot solve the problem of parameter space distortion at the root, and may lose important gradient information.
[0024] The second one is LayerNorm, which is a method of normalizing the input data, that is, calculating the mean and variance on the feature dimension of each sample, and then performing a normalization operation on all features of this sample. However, this method will change the probability distribution characteristics of the large language model, which will affect the accuracy of model training.
[0025] Therefore, in the embodiments of this specification, an improved method for training large language models is proposed, that is, by using the parameters of the normalization process of the output layer to map the input vector of the output layer. Thus, the numerical range of the output result of the output layer can be effectively controlled, which helps to improve the stability of the training of large language models.
[0026] The above improved solution can be applied to the training device for centralized training. Preferably, this improved solution can also be incorporated into the distributed training framework to further improve the training effect.
[0027] Figure 1 The schematic diagram of the distributed training framework in an example of this specification is shown. Figure 1 In this, the distributed training framework includes N GPUs, and the N GPUs can interact based on communication modes such as all_to_all (cross-device full exchange), all gather (full aggregation), reduce scatter (reduction scattering), all reduce (global reduction), etc.
[0028] In Figure 1 Under the shown distributed training framework, the model parameters of the large language model will be split onto N GPUs for parallel training. In other words, each GPU is responsible for training a part of the parameters of the large language model. In one example, the part of the parameters corresponding to each GPU includes the parameter shards corresponding to each network layer of the large language model.
[0029] Taking any network layer as an example, the original parameter matrix corresponding to this network layer can be sliced into multiple small matrices along the row direction. The number of columns of each small matrix is the same as that of the original parameter matrix. In this way, multiple parameter shards (i.e., small matrices) of this network layer can be obtained. Of course, the original parameter matrix corresponding to this network layer can also be sliced into multiple small matrices along the column direction. The number of rows of each small matrix is the same as that of the original parameter matrix. In this way, multiple parameter shards (i.e., small matrices) of this network layer can be obtained.
[0030] In addition, each GPU also holds several batches (also called local batches) among the multiple batches obtained by dividing the entire sample set (also called the global batch) used for training the large language model. For each local batch, it can be further divided into multiple micro-batches, and then these micro-batches are processed sequentially in order. For example, assuming the number of GPUs is 4 and the global batch size is 32, then the local batch size of each GPU is 8. If the local batch is further divided into 2 micro-batches, then the size of each micro-batch is 4.
[0031] It should be noted that due to the large number of parameters of the large language model, a large amount of memory is required to store intermediate results and gradients during the training process. By dividing the local batch into micro-batches, the GPU can process only a small part of the data each time, thus significantly reducing the memory usage.
[0032] In summary, in Figure 1 the distributed training framework shown, each GPU can calculate gradients independently, that is, each GPU can parallelly train the partial parameters of the large language model it holds, thereby accelerating the training speed of the large language model.
[0033] Figure 2 shows a flowchart of a method for training a large language model according to an embodiment of the present specification. This method can be executed by a training device for centralized training. Or, in a distributed training scenario, this method can be executed by Figure 1 any one of the GPUs in (for ease of description, called the first GPU). It should be noted that this method can include multiple rounds of iteration, Figure 2 shows the method steps included in the t-th (t is a positive integer) round of iteration. It can be understood that by repeatedly executing the steps shown therein, multi-round iterative updates of the parameters of the large language model (in a distributed training scenario, the parameter shards held by the first GPU) can be achieved. As Figure 2 shown, this method can include the following steps:
[0034] Step S202, obtain a training sample set including an input part and an output part.
[0035] Among them, the training sample set here can refer to the samples of a batch (hereinafter also referred to as the target batch) corresponding to the t-th round of iteration. In practice, the samples of a batch can be divided into multiple micro-batches, and the large language model processes the samples of one micro-batch each time.
[0036] The input part of any training sample in the above training sample set can include a piece of text. For example, it can be a sentence, an article, or a question. Thus, the input part can also be referred to as the input text. Specifically in the text generation task, the input text can be a piece of prompt text. For example, "Please describe the scenery in spring."
[0037] The output part corresponds to the input part and is the result expected to be output by the model. Specifically, it can be a correct answer, a complete piece of text, a classification label, etc. For example, for the above input text: "Please describe the scenery in spring.", the output part may be a detailed text describing the scenery in spring. Of course, the output part may also be: "Describe the scenery in spring?" etc. This specification does not limit this.
[0038] Step S204, input each input text in the training sample set into the large language model.
[0039] Among them, the large language model here can include an input layer, a hidden layer, and an output layer.
[0040] Furthermore, the input layer can include a word embedding layer and a position encoding layer. Among them, the word embedding layer is used to convert each word in the input text into a corresponding embedding vector. The position encoding layer is used to add position information to the embedding vectors of each word.
[0041] The hidden layer can include an encoder layer and / or a decoder layer. Further, the encoder layer can include a multi-head self-attention layer, a feed-forward neural network layer, and a layer normalization layer. Among them, the multi-head self-attention layer is used to calculate the degree of association between each word in the input text and other words, so as to capture the long-distance dependencies in the long text. The feed-forward neural network layer is used to perform a non-linear transformation on the output of the multi-head self-attention layer to further extract features. The layer normalization layer is used to normalize the features of each input text.
[0042] The decoder layer may include a masked multi-head self-attention layer, an encoder-decoder attention layer, a feed-forward neural network layer, and a layer normalization layer. Among them, the masked multi-head self-attention layer is similar to the multi-head self-attention layer. However, when generating text, in order to ensure the autoregressive property of the model, subsequent words need to be masked to prevent the model from seeing subsequent information. When the hidden layer includes both the encoder layer and the decoder layer, the encoder-decoder attention layer is used to fuse the output information of the encoder layer into the decoder layer, enabling the decoder layer to utilize the context information extracted by the encoder layer for text generation. The feed-forward neural network layer and the layer normalization layer have the same functions and principles as the corresponding layers in the encoder layer, and are used to perform non-linear transformation and normalization processing on the intermediate output of the decoder layer.
[0043] The above output layer is used to map the output of the hidden layer to a preset vocabulary space, thereby calculating a score for each token in the preset vocabulary space.
[0044] Of course, in practice, the above output layer may also include a softmax layer, which is used to convert the scores of each token into a probability distribution, such that the sum of the probabilities of all tokens is 1, facilitating the model to select the token with the highest probability as the generation result.
[0045] For any first input text, after inputting it into the large language model, the first input text is first processed through the parameters of the input layer and the hidden layer of the large language model (in a distributed training scenario, the parameter shards held by the first GPU for the input layer and the hidden layer), and then the corresponding vector representation (i.e., the output of the hidden layer) is obtained, which includes the representation vectors of each word in the first input text.
[0046] Then, the target parameter matrix of the output layer can be obtained, and through this target parameter matrix, the vector representation of the first input text (i.e., the input vector of the output layer) is mapped to obtain the output result mapped to the preset vocabulary space.
[0047] Here, different from the processing of the output layer in the conventional scheme, the target parameter matrix here is obtained by normalizing the original parameter matrix of the output layer obtained in the previous batch (i.e., in the (t - 1)-th round of iteration), so as to effectively control the numerical range of the output result of the output layer. If the current batch (i.e., the target batch) is the first batch of training, then the original parameter matrix here can refer to the initialized parameter matrix.
[0048] Specifically, since the output layer needs to implement the mapping from the output of the hidden layer to the preset vocabulary space, generally, each row in the original parameter matrix corresponds to each token in the preset vocabulary space. Correspondingly, the original parameter matrix can be normalized row by row, that is, the normalization process of the elements in each row is performed separately.
[0049] In a specific example, the above normalization process can be, for example, L2 norm normalization. Specifically, the formula for L2 norm normalization is as follows:
[0050]
[0051] In Formula 1, w 1 、w i 、…w n are the respective elements in any row of the original parameter matrix, and ‖W i ‖ 2 is the L2 norm of the row vector of this row, which is obtained by first calculating the sum of the squares of the elements in this row vector and then taking the square root.
[0052] Of course, in practice, various deformations can also be made to the above Formula 1. For example, ‖W‖ 2 can be replaced by other order norms or metrics, or replaced by learnable parameters, that is, more accurate values are gradually determined during the iterative training of the large language model. This specification does not make any limitations in this regard.
[0053] In a distributed training scenario, the first GPU only holds a parameter shard composed of some matrix elements in the original parameter matrix of the output layer, hereinafter referred to as the first parameter shard. As mentioned above, the output layer needs to implement the mapping from the output of the hidden layer (i.e., the input vector of the output layer) to the preset vocabulary space, so each row in the original parameter matrix corresponds to each token in the preset vocabulary space. In order to ensure that each GPU can map the input vector of the output layer to the entire vocabulary space, therefore, the column slicing method is usually used to slice the parameter matrix corresponding to the output layer. That is to say, the above first parameter shard is a small matrix obtained by slicing the original parameter matrix of the output layer along the column direction.
[0054] Since the parameter shards of each GPU are obtained by slicing the original parameter matrix along the column direction, and the normalization of the original parameter matrix needs to be performed row by row globally, before normalizing the original parameter matrix, the first GPU can collect other parameter shards from other GPUs respectively based on the cross-device full exchange communication mode (all_to_all communication mode), and then restore the original parameter matrix of the output layer based on each parameter shard. Then, according to the aforementioned normalization method, the original parameter matrix is normalized row by row to obtain a normalized matrix. The matrix part corresponding to the first parameter shard in the normalized matrix is used as the target parameter matrix for the first GPU.
[0055] In one embodiment, while other GPUs provide corresponding parameter shards to the first GPU, they can also provide corresponding location information (e.g., column identifiers), so that the first GPU can restore the original parameter matrix of the output layer based on this location information.
[0056] Figure 3 The figure shows a schematic diagram of the normalization processing method in an example of this specification. Figure 3 In it, the size of the parameter shards held by each GPU is n*m. Among them, each column in the parameter shard held by GPU1 is the 1st to the mth column of the original parameter matrix of the output layer, and the number of rows is the same as that of the original parameter matrix, that is, the number of rows is n rows. Each column in the parameter shard held by GPU2 is the (m + 1)th to the 2mth column of the original parameter matrix of the output layer, and the number of rows is the same as that of the original parameter matrix, and so on. Taking the first GPU as GPU1 as an example above, after normalizing the parameter shard it holds, the following can be obtained Figure 3 The target parameter matrix shown on the right.
[0057] It should be understood that in the above all_to_all communication mode, while the first GPU collects parameter shards from other GPUs respectively, it will provide the parameter shard it holds (i.e., the above first parameter shard) to other GPUs, that is, the first GPU is both a data sender and a data receiver.
[0058] In this solution, by using the above all_to_all communication mode, not only can the training speed be accelerated, but also the video memory consumption can be reduced.
[0059] It should be noted that in some scenarios, the original parameter matrix of the output layer of the large language model may not be differentiable. For example, in the scenario of reinforcement learning from human feedback (RLHF), the large language model is usually optimized by reinforcement learning algorithms (such as PPO). When these algorithms update the parameters of the large language model, they will introduce sampling of reward signals and policy gradient estimation, rather than directly backpropagating the reward signals. This indirect optimization method may lead to non-differentiability of some parameter update paths, and thus the original parameter matrix may not be differentiable.
[0060] In the case where the original parameter matrix is not differentiable, additionally, a first transformation operation is performed on the original parameter matrix to obtain a transformed original parameter matrix that is differentiable. Then, the transformed original parameter matrix is normalized row by row to obtain the above target parameter matrix.
[0061] In one embodiment, in the case where the original parameter matrix is not differentiable, the properties of the identity matrix can be borrowed to perform a first transformation operation on the original parameter matrix to make it differentiable.
[0062] In a more specific embodiment, the code statements corresponding to the above first transformation operation may be as follows:
[0063] unnormed_head = self.lm_head(self.input_matrix).transpose(1, 0);
[0064] Wherein, the above self.input_matrix is an identity matrix, the number of rows and columns of which is the same as the dimension of the input vector of the output layer, self.lm_head() is the processing function corresponding to the output layer, and transpose() is the dimension transposition function.
[0065] It should be noted that by executing the above code, the transformed original parameter matrix with differentiability can be obtained, and each row therein corresponds to each token in the preset vocabulary space.
[0066] It should be understood that in this solution, the target parameter matrices for processing each micro-batch divided from a batch are the same. Therefore, when processing the first micro-batch of each batch, normalization processing can be performed to obtain the above target parameter matrix, and then it is stored in the target storage space for use when processing subsequent micro-batches.
[0067] That is, when processing subsequent micro-batches (non-first micro-batches) of the same batch, the target parameter matrix can be directly read from the target storage space, which can improve the training speed of the large language model and save computing resources.
[0068] Finally, the output result corresponding to the above first input text may refer to the score of each token predicted for each position of the output part corresponding to the first input text by the target parameter matrix. Similarly to the mapping process of the first input text, each input text in the training sample set can be mapped through the target parameter matrix, and then the output results corresponding to each input text can be obtained.
[0069] Return to Figure 2 in Figure 2 The following steps may further be included:
[0070] Step S206, after obtaining the output results corresponding to each input text in each micro-batch included in the target batch, determining the target parameter gradient, and updating the original parameter matrix of the output layer using the target parameter gradient.
[0071] Specifically, first, through the softmax function, the scores of each token at each position in the output part predicted based on the first input text can be converted into a probability distribution, such that the sum of the probabilities of all tokens at each position is 1. In this way, the respective probability distributions corresponding to the first input text can be obtained. Subsequently, based on the respective probability distributions corresponding to the first input text and the respective true labels corresponding to the first input text, the classification loss corresponding to the first input text can be calculated based on the cross-entropy loss function. Among them, a single true label corresponding to the first input text is determined based on the true word at the corresponding position in the output part corresponding to the first input text.
[0072] Of course, in practice, the above cross-entropy loss function can also be replaced with a mean squared error loss function, etc., and this specification does not make a limitation in this regard.
[0073] After calculating the classification loss corresponding to a mini-batch (obtained by synthesizing the classification losses corresponding to the respective input texts in the mini-batch), the derivative of the target parameter matrix can be calculated based on this classification loss, and then the parameter gradient of the output layer corresponding to this mini-batch can be obtained.
[0074] It should be understood that the above parameter gradient is the gradient used to update the original parameter matrix of the output layer. In addition to calculating this parameter gradient, in practice, after calculating the classification loss corresponding to a mini-batch, the input gradient for backpropagation to the previous layer can also be calculated, which is specifically obtained by taking the derivative of the input vector of the output layer based on the classification loss. For this input gradient, it needs to be rotated back to the previous layer of the output layer to calculate the parameter gradient corresponding to the previous layer, and further calculate the input gradient rotated back to the layer before the previous layer, and so on, until reaching the first layer. In this way, the parameter gradients of each network layer corresponding to a mini-batch can be obtained. The solution of the embodiment of this specification focuses on the improvement of the output layer, and the gradient calculation and parameter update processes of other network layers are not elaborated in detail here.
[0075] It should be noted that since in the case where the original parameter matrix is not differentiable, the above target parameter matrix is obtained by first performing a first transformation operation on the original parameter matrix and then performing a row normalization process, therefore, for the parameter gradient obtained based on the target parameter matrix, a first operation corresponding to the first transformation operation (described later) needs to be performed on it to correct this parameter gradient.
[0076] In one embodiment, after obtaining the output results of each input text in an arbitrary micro-batch, the parameter gradients corresponding to the micro-batch are calculated; then, the first operation corresponding to the first transformation operation is performed on the parameter gradients corresponding to the micro-batch to obtain the corrected gradients corresponding to the micro-batch. After processing the current batch (i.e., the target batch), the corrected gradients corresponding to each micro-batch in the current batch are accumulated, and the target parameter gradients are obtained based on the accumulated results. Thus, the original parameter matrix of the output layer is updated using the target parameter gradients.
[0077] In another embodiment, to save computing resources, after obtaining the parameter gradients corresponding to a micro-batch, the first operation is not performed immediately, but only parameter gradient accumulation is carried out, that is, the above parameter gradients are accumulated with the parameter gradients corresponding to the previous micro-batches. After the last micro-batch of the current batch (i.e., the target batch) ends, the above first operation is performed on the accumulated parameter gradients to obtain the target parameter gradients.
[0078] Taking the case where the first transformation operation is performed by referring to the properties of the identity matrix as an example, the first operation here can specifically include: multiplying the accumulated parameter gradients by the identity matrix to obtain the target parameter gradients. In one example, the number of rows and columns of the identity matrix here is the same as the dimension of the input vector of the output layer.
[0079] It should be noted that in the solution of this embodiment, after processing each micro-batch, the first operation is not directly performed on the parameter gradients corresponding to the micro-batch, but on the accumulated parameter gradients of a batch, which can save computing resources and help improve the training efficiency of the large language model.
[0080] After obtaining the above target parameter gradients, the product between the preset learning rate and the target parameter gradients can be calculated, and the original parameter matrix is updated to the difference between it and the product, so that the original parameter matrix obtained in the t-th round of iterative training is obtained (in a distributed training scenario, that is, the parameter shards of the output layer obtained in the t-th round of iterative training are obtained).
[0081] It should be understood that the above is only a parameter update method. In practice, information such as the momentum of the target parameter gradients can also be combined to update the original parameter matrix, and this specification does not limit this.
[0082] For other network layers, the corresponding parameters can be directly updated based on the parameter gradients accumulated for each corresponding micro-batch, so that the parameters of other network layers obtained in the t-th round of iterative training can be obtained.
[0083] Thus, the t-th round of training of the parameters of each network layer of the large language model is completed.
[0084] It should be noted that after the original parameter matrix in the output layer is updated, the target parameter matrix stored in the target storage space can be deleted, so as to store a new target parameter matrix determined based on the original parameter matrix obtained in the t-th round of training in the next round of iteration, thereby realizing memory reuse and significantly reducing the occupation of memory space.
[0085] It should be understood that after training based on each batch, that is, after multiple rounds of iteration, the parameters of each network layer of the final training can be obtained. In a distributed training scenario, the first GPU obtains the parameter shards of each network layer of the final training. Similarly, each of the other GPUs can also obtain the parameters of each network layer of its final training, thus obtaining the large language model of the final training.
[0086] In summary, the method for training a large language model provided in the embodiments of this specification can effectively control the numerical range of the output result of the output layer while ensuring the continuity of its parameter space by normalizing the parameters of the output layer, which helps to improve the stability of the training of the large language model. In addition, this solution can be applied to a distributed training scenario, where, in the output layer, each GPU collects the parameter shards of other GPUs based on the all_to_all communication mode, which can avoid the problems of large data transmission volume and slow transmission speed brought by the allgather communication mode. Finally, this solution releases the memory space used to store the target parameter matrix after the training of one batch is completed, which can significantly reduce the occupation of memory space.
[0087] Corresponding to the method for training a large language model, an embodiment of this specification also provides an apparatus for training a large language model, and the large language model includes an output layer. As Figure 4 shown, the apparatus includes:
[0088] An obtaining unit 402, configured to obtain a target parameter matrix of the output layer, which is obtained by normalizing the original parameter matrix obtained by the output layer in the previous batch of training.
[0089] A mapping unit 404, configured to perform a mapping process on the input vector of the output layer through the target parameter matrix to obtain an output result mapped to a preset vocabulary space, and the input vector corresponds to the input text.
[0090] A determining unit 406, configured to determine a target parameter gradient after obtaining the output results corresponding to each input text in each micro-batch included in the current batch, and update the original parameter matrix by using the target parameter gradient.
[0091] In one embodiment, the obtaining unit 402 includes:
[0092] An obtaining sub-module 4022, configured to obtain the original parameter matrix;
[0093] The transformation sub-module 4024 is configured to perform a first transformation operation on the original parameter matrix to obtain a transformed original parameter matrix when the original parameter matrix is not differentiable, where each row corresponds to each token in a preset vocabulary space;
[0094] The normalization sub-module 4026 is configured to perform a normalization process on the transformed original parameter matrix row by row to obtain a target parameter matrix.
[0095] In one embodiment, the determining unit 406 is specifically configured to:
[0096] Obtain an accumulated parameter gradient, which is accumulated from the parameter gradients corresponding to each mini-batch. The parameter gradient corresponding to a single mini-batch is calculated based on the output results of each input text in the mini-batch and the target parameter matrix;
[0097] Perform a first operation corresponding to the first transformation operation on the accumulated parameter gradient to obtain a target parameter gradient.
[0098] In one embodiment, the determining unit 406 is further specifically configured to:
[0099] Multiply the accumulated parameter gradient by the identity matrix to obtain a target parameter gradient.
[0100] In another embodiment, the determining unit 406 is specifically configured to:
[0101] After obtaining the output results of each input text in any mini-batch, calculate the parameter gradient corresponding to the mini-batch;
[0102] Perform a first operation corresponding to the first transformation operation on the parameter gradient corresponding to the mini-batch to obtain a corrected gradient corresponding to the mini-batch;
[0103] Accumulate the corrected gradients corresponding to each mini-batch in the current batch, and obtain a target parameter gradient based on the accumulation result.
[0104] In one embodiment, the above normalization is L2 norm normalization.
[0105] In one embodiment, the above input text belongs to the first mini-batch in the current batch, and the apparatus further includes:
[0106] The storage unit 408 is configured to perform a normalization process on the original parameter matrix and store the obtained target parameter matrix in a target storage space.
[0107] In one embodiment, the above input text belongs to a non-first mini-batch in the current batch, and the obtaining unit 402 is specifically configured to:
[0108] Read the target parameter matrix from the target storage space, where the target parameter matrix is stored in the target storage space when the first micro-batch of the current batch is deposited.
[0109] In one embodiment, the above device is arranged on any first GPU in the GPU cluster, which holds a first parameter shard composed of some matrix elements in the original parameter matrix; each row in the original parameter matrix corresponds to each token in the preset vocabulary space.
[0110] The obtaining unit 402 is specifically configured to:
[0111] Receive other parameter shards from other GPUs in the GPU cluster respectively, and restore the original parameter matrix based on each parameter shard.
[0112] Perform normalization processing on the original parameter matrix row by row to obtain a normalized matrix, and determine the part corresponding to the first parameter shard in the normalized matrix as the target parameter matrix.
[0113] In one embodiment, the device further includes:
[0114] A sending unit 410, configured to provide the first parameter shard to other GPUs.
[0115] The functions of the functional units of the device in the above embodiments of this specification can be implemented by the steps of the above method embodiments. Therefore, the specific working process of the device provided in an embodiment of this specification will not be repeated here.
[0116] The device for training a large language model provided in an embodiment of this specification can improve the stability of training the large language model.
[0117] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in combination with Figure 2 as described.
[0118] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor. An executable code is stored in the memory, and when the processor executes the executable code, the method described in combination with Figure 2 as described is implemented.
[0119] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the medium or device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0120] The specific embodiments of the present specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0121] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present specification. It should be understood that the above description is only the specific embodiments of the present specification and is not used to limit the protection scope of the present specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present specification shall be included within the protection scope of the present specification.
Claims
1. A method for training a large language model, the large language model comprising an output layer; the method comprising: Obtaining a target parameter matrix of the output layer, which is obtained by normalizing an original parameter matrix of the output layer obtained in a previous batch of training; The input vector of the output layer is mapped by the target parameter matrix to obtain an output result mapped to a preset vocabulary space; the input vector corresponds to the input text; After obtaining the output results corresponding to each input text in each micro-batch included in the current batch, the target parameter gradient is determined, and the original parameter matrix is updated using the target parameter gradient.
2. The method according to claim 1, wherein: The step of obtaining a target parameter matrix of the output layer comprises: Obtaining the original parameter matrix; In the case where the original parameter matrix is not differentiable, performing a first transformation operation on the original parameter matrix to obtain a transformed original parameter matrix, wherein each row corresponds to each word in a preset word table space; The transformed original parameter matrix is normalized row by row to obtain the target parameter matrix.
3. The method according to claim 2, wherein: Determining the target parameter gradient includes: Obtaining a cumulative parameter gradient, which is accumulated from the parameter gradients corresponding to the micro-batches, wherein the parameter gradient corresponding to a single micro-batch is calculated based on the output results of each input text in the micro-batch and the target parameter matrix; A first operation corresponding to the first transform operation is performed on the accumulated parameter gradient to obtain the target parameter gradient.
4. The method according to claim 3, wherein: The performing a first operation corresponding to the first transformation operation on the accumulated parameter gradient includes: The accumulated parameter gradient is multiplied by the identity matrix to obtain the target parameter gradient.
5. The method according to claim 2, wherein: Determining the target parameter gradient includes: After obtaining the output results of each input text in any micro-batch, the parameter gradient corresponding to the micro-batch is calculated; Performing a first operation corresponding to the first transformation operation on the parameter gradient corresponding to the micro-batch to obtain a correction gradient corresponding to the micro-batch; The correction gradients corresponding to each micro-batch in the current batch are accumulated, and the target parameter gradient is obtained based on the accumulation result.
6. The method according to claim 1 or 2, wherein: The normalization is L2 norm normalization.
7. The method according to claim 1, wherein: The input text belongs to the first micro-batch in the current batch, and obtaining the target parameter matrix of the output layer includes: The original parameter matrix is normalized, and the obtained target parameter matrix is stored in a target storage space.
8. The method according to claim 1, wherein: The input text belongs to a non-first micro-batch in the current batch, and obtaining the target parameter matrix of the output layer includes: The target parameter matrix is read from a target storage space, and the target parameter matrix is stored in the target storage space in a first micro-batch of the current batch.
9. The method according to claim 1, wherein: The method is executed by any first GPU in the GPU cluster, which holds a first parameter slice consisting of some matrix elements in the original parameter matrix; Each row in the original parameter matrix corresponds to each word in the preset word table space; The step of obtaining a target parameter matrix of the output layer comprises: Receiving other parameter slices from other GPUs in the GPU cluster respectively, and restoring the original parameter matrix based on each parameter slice; The original parameter matrix is normalized row by row to obtain a normalized matrix, and a portion of the normalized matrix corresponding to the first parameter slice is determined as the target parameter matrix.
10. The method according to claim 9, further comprising: Send the first parameter slice to each of the other GPUs.
11. A device for training a large language model, the large language model comprising an output layer; the device comprising: An acquisition unit, used for acquiring a target parameter matrix of the output layer, which is obtained by normalizing an original parameter matrix of the output layer obtained in a previous batch of training; A mapping unit, used to map the input vector of the output layer through the target parameter matrix to obtain an output result mapped to a preset vocabulary space; The input vector corresponds to the input text; The determination unit is used to determine the target parameter gradient after obtaining the output results corresponding to each input text in each micro-batch included in the current batch, and update the original parameter matrix using the target parameter gradient.
12. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 10.
13. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 10 is implemented.