Method and device for training large language model
Through the use of sliding window strategies and dynamic reference values, abnormal situations in the training process of large language models are monitored and handled in real time, training exceptions are solved and training efficiency and stability are improved.
Patent Information
- Application Number
- CN202510213820.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
During the training process of large language models, training exceptions may occur due to data noise, model complexity or improper settings of optimizer parameters, resulting in the update direction of the model parameter deviating from the normal trajectory, reducing training efficiency or causing the model to crash.
The process data statistical characteristics during the training process are tracked through the sliding window strategy, and the abnormalities of the training round are monitored in real time using dynamic benchmark values. When the difference between the process data of the target training round and the reference value exceeds the preset threshold, it is determined that the round is an abnormal training round and performs corresponding processing, including skipping or adjusting the hyperparameters to reduce the impact of the abnormality.
By monitoring and handling abnormal situations during the training process in real time, the efficiency and stability of training of large language models are improved, model crashes are avoided, and training effects are enhanced.
Smart Images

Figure CN120046685A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of machine learning, and in particular, to methods and devices for training large language models. Background Art
[0002] In recent years, neural networks have made significant progress in multiple fields, demonstrating powerful data analysis and processing capabilities. With the expansion of the scale of neural networks and the improvement of hardware computing capabilities, pre-trained models with even larger scales have emerged, promoting the development of large language models (LLMs). Large language models are trained with large-scale data and have powerful text understanding and generation capabilities, and are widely used in multiple fields.
[0003] However, during the model training process, due to reasons such as data noise, model complexity, or improper optimizer parameter settings, anomalies sometimes occur during training. These anomalies can cause the model parameter update direction to deviate from the normal trajectory, reduce the training efficiency, and even lead to model collapse. Therefore, a method is needed to address the anomalies that occur during the model training process so that the model can complete training smoothly. Summary of the Invention
[0004] One or more embodiments of this specification describe methods and devices for training large language models to jointly utilize the text information and relationship information in semi-structured data to assist the model in answering relevant questions to generate answers.
[0005] In a first aspect, a method for training a large language model is provided, including:
[0006] By inputting the training samples of the target batch into the large language model, determining the process data of the target training round, the training samples including text data, and the process data including the training loss value or the gradient value of each parameter;
[0007] Obtaining a reference value obtained by statistically analyzing the process data of N consecutive training rounds before the target training round;
[0008] When the target difference between the process data of the target training round and the reference value exceeds a preset first threshold, determining the target training round as an abnormal training round;
[0009] Performing target processing on the abnormal training round; the target processing includes skipping the abnormal training round or adjusting the hyperparameters in the abnormal training round to reduce the impact of the abnormal training round.
[0010] In some possible embodiments, the process data includes gradient values of various parameters; the target difference includes the difference between the target gradient value of any parameter and the corresponding reference value of the parameter.
[0011] In some possible embodiments, the target difference is the ratio of the process data of the target training round to the reference value.
[0012] In some possible embodiments, the target processing is to skip the abnormal training round; the method further includes:
[0013] Put the training samples of the target batch back into the training sample pool for resampling.
[0014] In some possible embodiments, the target processing is to adjust the hyperparameters in the abnormal training round; the hyperparameters include the learning rate; the target processing specifically includes:
[0015] According to the target difference, reduce the value of the learning rate, and update the large language model according to the adjusted learning rate.
[0016] In some possible embodiments, the target processing is to adjust the hyperparameters in the abnormal training round; the hyperparameters include the learning rate; the target processing specifically includes:
[0017] Reduce the value of the learning rate to a first preset value, and update the large language model according to the adjusted learning rate.
[0018] In some possible embodiments, it further includes:
[0019] After the target processing, put the training samples of the target batch back into the training sample pool for resampling.
[0020] In some possible embodiments, putting the training samples of the target batch back into the training sample pool includes:
[0021] Mark each training sample in the target batch as an abnormal training sample, and multiply the sampling probability of each abnormal training sample by a preset first magnification factor; the first magnification factor is less than 1;
[0022] Put each abnormal training sample back into the training sample pool so that it is sampled with its respective sampling probability in subsequent training rounds.
[0023] In some possible embodiments, the process data includes gradient values of various parameters; the method further includes:
[0024] When the target difference does not exceed the first threshold, determine the average gradient value of each parameter according to the gradient values of each parameter in the N training rounds;
[0025] Determine the norm and covariance matrix determined according to the average gradient values of the respective parameters, and determine the gradient noise scale value;
[0026] When the gradient noise scale value is greater than a preset second threshold, determine the target training round as an abnormal training round.
[0027] In some possible implementation manners, the average gradient value of each parameter is determined according to the exponentially weighted moving average.
[0028] In some possible implementation manners, the target processing is to skip the abnormal training round; the method further includes:
[0029] Increase the number of training samples used in each training round after the target training round.
[0030] In some possible implementation manners, it further includes:
[0031] When there are M consecutive training rounds determined to be abnormal training rounds, terminate the training process of the large language model.
[0032] In a second aspect, there is provided an apparatus for training a large language model, including:
[0033] A process data determination unit configured to determine the process data of a target training round by inputting a target batch of training samples into the large language model, where the training samples include text data, and the process data includes a training loss value or gradient values of each parameter;
[0034] An acquisition unit configured to acquire a reference value obtained by statistically processing the process data of N consecutive training rounds before the target training round;
[0035] A first determination unit configured to determine the target training round as an abnormal training round when the target difference between the process data of the target training round and the reference value exceeds a preset first threshold;
[0036] An abnormal processing unit configured to perform target processing on the abnormal training round; the target processing includes skipping the abnormal training round or adjusting hyperparameters in the abnormal training round to reduce the impact of the abnormal training round.
[0037] In a third aspect, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.
[0038] In a fourth aspect, a computing device is provided, including a memory and a processor. The memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.
[0039] The method and device for training a large language model proposed in the embodiments of this specification track the statistical characteristics of process data during the training of the large language model through a sliding window strategy, and use the statistical characteristics as a reference value. Then, according to the difference between the corresponding process data in subsequent rounds and the reference value, when the difference is too large, it is considered that the training in this round is abnormal. Specifically, the method compares the process data of the current training round with the reference value determined by the previous N consecutive training rounds before this training round, and determines whether the current round is an abnormal training round. When an abnormality exists, corresponding processing is performed on this training round. The processing methods include skipping this training round or adjusting the values of the hyperparameters in this training round to reduce the impact of this abnormal training round.
[0040] The method and device for training a large language model proposed in the embodiments of this specification use a dynamic reference value, which can adapt to the different degrees of fluctuations of the process data generated during the training process in different training stages and make corresponding processing in real time. And through a lightweight statistical calculation method, real-time monitoring of the training process can be realized, and corresponding processing can be automatically performed without complex calculations or manual intervention, improving the automation degree of the overall training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only the multiple embodiments disclosed in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 Schematic diagram of an implementation scenario of a method for training a large language model according to an embodiment;
[0043] Figure 2 Flowchart of a method for training a large language model according to an embodiment;
[0044] Figure 3 Schematic block diagram of a device for training a large language model according to an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The following describes the solutions provided in this specification with reference to the drawings.
[0046] During the training process of large language models, due to reasons such as training data noise, the complexity of the model itself, and improper optimizer parameter settings, training anomalies sometimes occur, including spike anomalies in training loss values and / or parameter gradient values. Among them, the spike in the training loss value is mainly manifested as the abnormal phenomenon that the function value of the loss function suddenly increases sharply in a short period during training; the spike anomaly in the parameter gradient value is mainly manifested as the abnormal phenomenon that the gradient of the trainable parameters in the model undergoes a drastic change in magnitude or direction in a short period during the backpropagation of training. These anomalies will affect the convergence and stability of model training. If not effectively controlled, these anomalies may cause the direction of model parameter update to deviate from the normal trajectory, resulting in a decrease in training efficiency at best, making the model converge slowly or even stagnate, and at worst, causing the model to collapse, seriously affecting the final training effect.
[0047] For the spike anomaly problem, the coping methods in related technologies include using a fixed learning rate or static gradient clipping, etc. However, most of these methods are static methods, lacking adaptability and being difficult to dynamically adjust to cope with the complex situations in different training stages. For example, a fixed learning rate may lead to slow convergence in the early stage and excessive oscillation in the later stage; although static gradient clipping can prevent extreme gradient values, it may also truncate valuable information and affect the final performance of the model.
[0048] Generally speaking, the entire training process of large language models will include multiple training rounds. In any training round, the training samples are input into the large language model for forward propagation to obtain the output results corresponding to each training sample. Then, the training loss is determined based on the output results, and backpropagation is performed in the large language model according to the training loss. During backpropagation, the partial derivative of the training loss with respect to each trainable parameter (such as weights, bias values, etc., which will also be simply referred to as "parameters" below) in the large language model is calculated to obtain the gradient value corresponding to each parameter. Then, the values of each parameter are updated according to the gradient values of each parameter, as shown in formula (1).
[0049]
[0050] Among them, θ represents the trainable parameters in the large language model, ← is the assignment symbol, α is the learning rate, J(θ) is the loss function, is the gradient value of the loss function J(θ) with respect to the parameter θ. It can be seen that when the loss J(θ) or the gradient has a spike anomaly, it will have a direct impact on the update of the parameter θ.
[0051] To overcome the above problems, an embodiment of this specification proposes a method and apparatus for training a large language model. The method uses a dynamic reference value and adjusts the training process in real time according to the monitoring results to improve the efficiency and accuracy of large language model training.
[0052] Figure 1 The following figure shows a schematic diagram of an implementation scenario of a method for training a large language model according to an embodiment. As Figure 1 shown, first, process data generated during the past consecutive N training rounds of the training process is obtained, including training loss values or gradient values of each parameter, and numerical statistical processing is performed on these N sets of process data to obtain a reference value. Numerical statistics can be, for example, the average, median, maximum, etc. On the other hand, the training samples of the current batch are input into the large language model to obtain the process data generated during the training of the current training round. It can be understood that the process data used in the current round should be of the same type as the process data used when calculating the reference value, that is, both are training loss values or both are gradient values of parameters. Next, the process data of the current round is compared with the reference value. When the difference between the process data and the reference value is too large, for example, the difference or ratio between the two exceeds a preset threshold, the current round is determined as an abnormal training round.
[0053] After determining a training round as an abnormal training round, corresponding processing needs to be performed on it to eliminate or alleviate the overall impact of the abnormality on model training. The processing methods include skipping the abnormal training round or adjusting the hyperparameters in the abnormal training round to reduce the impact of the abnormal training round. When a certain training round is determined to be an abnormal training round, it means that the sample quality of the training samples used in this round is relatively poor. Therefore, corresponding processing measures need to be taken for processing.
[0054] When choosing to skip the abnormal training round, the training of the next round is directly carried out, and the loss function value and parameter gradient value of this round will not affect the value of the model parameters. Optionally, since skipping a round is equivalent to not using the training samples of this batch, in order to effectively utilize the training samples, these training samples can be marked as abnormal training samples and then put back into the sample pool so that they can be resampled with a certain probability and used for training in subsequent training rounds.
[0055] When choosing to adjust the hyperparameters in the abnormal training round, the loss function value and parameter gradient value of this round will continue to be used to update the model parameters, but the impact of this abnormal round on model training will be reduced by adjusting the hyperparameters (such as the learning rate) used in the training process.
[0056] It should be noted that Figure 1It only shows the process of anomaly judgment and handling for one training round. In other embodiments, the processes of judgment and handling as shown in Figure 1 can be performed for each round during the training process. In this way, a sliding window is used to dynamically statistically analyze the training data of historical rounds and generate a dynamic baseline value, based on which a dynamic evaluation of the training process can be achieved.
[0057] The following describes the specific implementation steps of the above method for training a large language model in combination with specific embodiments.
[0058] Figure 2 A flowchart showing a method for training a large language model according to an embodiment, and the execution subject of the method can be any platform or server or device cluster with computing and processing capabilities, etc. As Figure 2 shown, the method at least includes: Step 202, determining the process data of the target training round by inputting a target batch of training samples into the large language model, where the training samples include text data, and the process data includes a training loss value or gradient values of each parameter; Step 204, obtaining a baseline value statistically calculated from the process data of N consecutive training rounds before the target training round; Step 206, when the target difference between the process data of the target training round and the baseline value exceeds a preset first threshold, determining the target training round as an abnormal training round; Step 208, performing target processing on the abnormal training round; the target processing includes skipping the abnormal training round or adjusting hyperparameters in the abnormal training round to reduce the impact of the abnormal training round.
[0059] The following describes the specific execution processes of the above steps.
[0060] First, in Step 202, the process data of the target training round is determined by inputting a target batch of training samples into the large language model, where the training samples include text data, and the process data includes a training loss value or gradient values of each parameter.
[0061] The target batch of training samples can be sampled from a training sample pool containing several training samples when using a non-full gradient descent method for training (i.e., not using all samples in the training sample pool in each round of training), such as using the Mini-batch Gradient Descent method to train the large language model. The target batch of training samples is used to train the large language model in the target training round.
[0062] The target training round can be any round (including the second round) after the second round in the multiple rounds of training for the large language model.
[0063] The process data can be intermediate process data generated during the training process, including training loss values or gradient values of various parameters.
[0064] The training samples include text data, which is used to guide the large language model to understand the semantic information in the text, such as input text data, or paired input text-output text data. In the case where the large language model supports multi-modal, in addition to text data, the training samples can also include data of other modalities, such as image data.
[0065] In addition, in step 204, a reference value obtained by statistically analyzing the process data of N consecutive training rounds before the target training round is acquired.
[0066] The size of the sliding window can be set to a first fixed value K. When the round number n of the target round is less than or equal to K, that is, the target round is a round in the starting stage, and the number of accumulated rounds passed is less than K rounds, then let N = n - 1, that is, use the process data of all training rounds before the target round for statistics to obtain the reference value; when the round number n of the target round is greater than K, then let N = K, that is, use the process data of K training rounds before the target round for statistics to obtain the reference value.
[0067] Multiple statistical metrics can be used for statistically analyzing the process data of N consecutive training rounds, such as average value, median, maximum value, etc., which are not limited here.
[0068] It can be understood that the process data of N consecutive training rounds should correspond to the type of process data determined in the target training round in step 202. That is, when the type of process data determined in step 202 is the training loss value, the corresponding reference value is the statistical calculation result of the training loss value, and the process data of N consecutive training rounds in step 204 is also the training loss value; when the type of process data determined in step 202 is the gradient value of each parameter, the corresponding reference value is the statistical result of the gradient value of each parameter, and the process data of N consecutive training rounds in step 204 is also the gradient value of each parameter.
[0069] By statistically analyzing the process data of each round in the sliding window, a reference value can be obtained. This reference value can be used to evaluate whether there is an abnormality in the current target training round.
[0070] Next, in step 206, when the target difference between the process data of the target training round and the reference value exceeds a preset first threshold, the target training round is determined as an abnormal training round.
[0071] Discussions can be carried out separately according to the specific content included in the process data. In one embodiment, the process data includes a training loss value, that is, when the training loss value of the target training round is too large relative to the reference value, the target training round is determined as an abnormal training round.
[0072] In another embodiment, the process data includes the gradient values of each parameter. Since a large language model usually includes multiple trainable parameters, the number of corresponding gradient values is also multiple. Therefore, the calculation of the target difference can be based on some or all of these gradient values.
[0073] In a more specific embodiment, the process data includes the gradient values of each parameter; the target difference is the difference between the first norm of the gradient vector composed of the gradient values of each parameter and the reference value determined by the statistical result of the norms of the gradient vectors in each round. Or, in another specific example, the target difference can be the vector distance between the gradient vector composed of the gradient values of each parameter in this round and the average gradient vector (as the reference value) in the previous N rounds.
[0074] In another more specific embodiment, the process data includes the gradient values of each parameter; the target difference includes the difference between the target gradient value of any parameter and the reference value corresponding to this parameter.
[0075] Multiple ways can be used to measure the target difference between the process data of the target training round and the reference value, such as the difference, ratio, etc. between the two, which are not limited here.
[0076] In one embodiment, the target difference is the ratio between the process data of the target training round and the reference value. When this ratio exceeds a preset first threshold, that is, when the value of the process data of the target training round is too large, the target training round is determined as an abnormal training round.
[0077] Then, in step 208, target processing is performed on the abnormal training round; the target processing includes skipping the abnormal training round, or adjusting the hyperparameters in the abnormal training round to reduce the impact of this abnormal training round.
[0078] When the target training round is determined as an abnormal training round in step 206, target processing will be performed on this training round to eliminate or reduce the impact of this abnormality on the overall training process.
[0079] In a possible implementation, the target processing is to skip the abnormal training round. When the process data is the training loss value, skipping the abnormal training round means that it is not necessary to backpropagate to calculate the gradient, but directly proceed to the next training round. When the process data is the gradient value, skipping the abnormal training round means skipping the process of updating the parameters of this training round, that is, the process shown in formula (1), and directly proceeding to the next training round.
[0080] In one embodiment, the skipped training rounds do not participate in the subsequent sliding window-based baseline value statistics.
[0081] When choosing to skip the abnormal training rounds, it is equivalent to not using the training samples of the target batch for actual training and model parameter update, which to some extent causes waste of training samples. Therefore, in one implementation, the method further includes step 210.
[0082] In step 210, the training samples of the target batch are returned to the training sample pool for resampling.
[0083] Resampling means resampling again, that is, enabling the training samples of the target batch to be resampled and participate in training in subsequent training rounds. The detailed process of resampling will be described later.
[0084] To cooperate with the implementation of resampling, in a more specific implementation, the process of sampling a batch of training samples from the training sample pool in each normal training round is a sampling process without replacement, that is, the sampled samples will be removed from the training sample pool.
[0085] In another possible implementation, the target processing is to adjust the hyperparameters in the abnormal training round to reduce the impact of this abnormal training round. Here, the hyperparameters can include the learning rate in formula (1), or the scaling factor set for the loss value, or other hyperparameters that affect parameter update. In one implementation, the abnormal training rounds processed by hyperparameter adjustment can participate in the subsequent sliding window-based baseline value statistics. Of course, it is also possible to set the sliding window-based baseline value statistics to not count any abnormal training rounds.
[0086] In one embodiment, the hyperparameters include the learning rate. Correspondingly, the target processing specifically includes: reducing the value of the learning rate according to the target difference, and updating the large language model according to the adjusted learning rate.
[0087] According to the target difference, reducing the value of the learning rate may specifically include reducing the specific value of the learning rate according to the adjustment range determined by the target difference. When the target difference is a ratio, reducing the specific value of the learning rate may include proportionally reducing the value of the learning rate according to this ratio; when the target difference is a difference, reducing the specific value of the learning rate may subtract this difference or a multiple of this difference from the specific value of the learning rate.
[0088] It should be noted that the value of the learning rate after reduction is only valid for the current target round. In the next training round, the original learning rate is still used for training.
[0089] In this embodiment, the learning rate is reduced in proportion or by a corresponding difference to mitigate the impact of this abnormal training round on model training.
[0090] In another possible embodiment, the target processing specifically includes:
[0091] Reduce the value of the learning rate to a first preset value, and update the large language model according to the adjusted learning rate.
[0092] In one example, the first preset value may be a value that is very small relative to the learning rate, such as 1 / 100 of the learning rate. That is, when this training round is an abnormal training round, a very small learning rate is used to complete the training process of this round.
[0093] It should be noted that the value of the learning rate after reduction is only valid for the current target round. In the next training round, the original learning rate is still used for training.
[0094] In this embodiment, since a very small learning rate is used, it is equivalent to the training samples of the target batch used in this round not being fully utilized, which also causes a certain degree of waste of training samples to some extent. Therefore, similar to the first embodiment, in this embodiment, the method further includes step 210.
[0095] In step 210, put the training samples of the target batch back into the training sample pool for resampling.
[0096] The above separately describes the specific implementation processes of two types of target processing. The following describes the detailed process of step 210.
[0097] In one embodiment, putting the training samples of the target batch back into the training sample pool in step 210 includes step 2102 and step 2104.
[0098] In step 2102, each training sample in the target batch is marked as an abnormal training sample, and the sampling probability of each abnormal training sample is multiplied by a preset first magnification factor; the first magnification factor is less than 1.
[0099] Initially, the sampling probabilities of all training samples in the training sample pool are equal, for example, all are 1. When some training samples are marked as abnormal training samples, the sampling probabilities of these abnormal training samples will be reduced. Specifically, the sampling probability can be multiplied by a preset first magnification factor less than 1. So that it is sampled with a smaller probability in subsequent training rounds and mixed with other normal samples to form a mini-batch sample for training the large language model. While reducing the impact of abnormal samples on the model, it can also effectively utilize each training sample.
[0100] In step 2104, the respective abnormal training samples are put back into the training sample pool so that they are sampled with their respective sampling probabilities in subsequent training rounds.
[0101] In one embodiment, an upper limit can also be set for the number of times of resampling of the training samples. Specifically, when the number of times a certain training sample is marked as an abnormal training sample exceeds a preset third threshold, then this training sample is removed from the sample training pool.
[0102] Through steps 202 to 210, the embodiments of this specification can adapt to the normal fluctuation range in different training stages through the dynamic threshold of statistical features based on a sliding window, avoiding the overfitting problem of a fixed threshold. At the same time, the embodiments of this specification propose multiple processing methods such as skipping abnormal training rounds, learning rate scaling, and resampling, avoiding the limitations of a single strategy and being able to flexibly handle different abnormal scenarios. And real-time detection is achieved through lightweight statistical calculations, without manual intervention or complex calculations. In addition, by recording abnormal samples and supporting limited retries, data waste is reduced while ensuring stability.
[0103] In some possible implementation manners, after making a determination using the reference value, the gradient noise scale value can also be combined to make a secondary determination for the training rounds that are not determined to be abnormal training rounds. In this implementation manner, the process data includes the gradient values of each parameter; the method further includes steps (a) to (c).
[0104] In step (a), when the target difference does not exceed the first threshold, the average gradient value of each parameter is determined according to the gradient values of each parameter in the N training rounds.
[0105] In one embodiment, the average gradient value of each parameter is determined according to the Exponential Moving Average (EMA). Using the exponential moving average can make the average gradient value smoother.
[0106] When the target training round is not identified as an abnormal training round in step 206, step (a) is entered to determine the average gradient value of each parameter in N training rounds.
[0107] In step (b), the gradient noise scale value is determined according to the norm and covariance matrix determined by the average gradient value of each parameter.
[0108] The calculation method of the gradient noise scale B can be as shown in formula (2):
[0109]
[0110] Where Σ is the covariance matrix of the gradients of each parameter, tr(Σ) is the trace of the covariance matrix, μ is the gradient vector determined by the average gradient value of each parameter, and ||μ|| is the norm of μ, for example, the second norm.
[0111] In step (c), when the gradient noise scale value is greater than a preset second threshold, the target training round is determined as an abnormal training round.
[0112] The gradient noise scale is used to measure the intensity of the gradient noise. When the value of the gradient noise scale is too large, it will affect the stability of the training process and slow down the convergence speed of the training. Therefore, when the gradient noise scale value is greater than a preset second threshold, the target training round is determined as an abnormal training round, and step 208 is entered.
[0113] In one embodiment, the target processing in step 208 is to skip the abnormal training round. The method further includes:
[0114] Step (d), increasing the number of training samples used in each training round after the target training round.
[0115] When the value of the gradient noise scale is too large, the number of training samples in each sample batch in subsequent training rounds, that is, the size of the mini batch, can be increased to improve the stability of the training.
[0116] The process of anomaly analysis and corresponding processing for a single training round is described above. In some possible implementation manners, the method further includes: step (e), when M consecutive training rounds are determined to be abnormal training rounds, terminating the training process of the large language model. That is, when multiple training rounds in a certain stage continuously show anomalies, it indicates that there are relatively serious problems in the training process. At this time, terminating the training process of the large language model can save unnecessary waste of computing resources and perform manual inspection on the settings of the training process to find specific problems and fix them.
[0117] According to an embodiment of another aspect, there is also provided an apparatus for training a large language model. Figure 3 A schematic block diagram of an apparatus for training a large language model according to an embodiment is shown. The apparatus can be deployed in any device, platform, or cluster of devices with computing and processing capabilities. As Figure 3 shown, the apparatus 300 includes:
[0118] A process data determination unit 302, configured to determine process data of a target training round by inputting a training sample of a target batch into the large language model. The training sample includes text data, and the process data includes a training loss value or gradient values of various parameters;
[0119] An acquisition unit 304, configured to acquire a reference value obtained by statistically analyzing process data of N consecutive training rounds before the target training round;
[0120] A first judgment unit 306, configured to determine the target training round as an abnormal training round when a target difference between the process data of the target training round and the reference value exceeds a preset first threshold;
[0121] An anomaly processing unit 308, configured to perform target processing on the abnormal training round; the target processing includes skipping the abnormal training round or adjusting hyperparameters in the abnormal training round to reduce the impact of the abnormal training round.
[0122] In some possible implementation manners, the apparatus 300 further includes:
[0123] A return unit, configured to return the training sample of the target batch to the training sample pool for resampling.
[0124] In some possible implementation manners, the process data includes gradient values of various parameters; the apparatus 300 further includes:
[0125] An average gradient value determination unit, configured to determine the average gradient value of each parameter according to the gradient values of each parameter in the N training rounds when the target difference does not exceed the first threshold;
[0126] A gradient noise scale determination unit, configured to determine a gradient noise scale value according to the norm and covariance matrix determined by the average gradient value of each parameter;
[0127] A second determination unit, configured to determine the target training round as an abnormal training round when the gradient noise scale value is greater than a preset second threshold.
[0128] In some possible implementation manners, the target processing is to skip the abnormal training round; the apparatus 300 further includes:
[0129] A quantity increase unit, configured to increase the quantity of training samples used in each training round after the target training round.
[0130] In some possible implementation manners, the apparatus 300 further includes:
[0131] A termination unit, configured to terminate the training process of the large language model when there are M consecutive training rounds determined as abnormal training rounds.
[0132] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method described in any of the above embodiments.
[0133] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor, wherein an executable code is stored in the memory, and when the processor executes the executable code, the method described in any of the above embodiments is implemented.
[0134] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0135] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0136] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0137] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0138] The specific embodiments described above further elaborate on the purpose, technical solution and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for training a large language model, comprising: Determining process data of a target training round by inputting a target batch of training samples into the large language model, wherein the training samples include text data, and the process data includes a training loss value or a gradient value of each parameter; Obtain a benchmark value obtained by statistically analyzing process data of N consecutive training rounds before the target training round; When a target difference between the process data of the target training round and the reference value exceeds a preset first threshold, determining the target training round as an abnormal training round; The abnormal training round is subjected to target processing; the target processing includes skipping the abnormal training round, or adjusting a hyperparameter in the abnormal training round to reduce the impact of the abnormal training round.
2. The method according to claim 1, wherein: The process data includes the gradient values of various parameters; the target difference includes the difference between the target gradient value of any parameter and the reference value corresponding to the parameter.
3. The method according to claim 1, wherein: The target difference is the ratio between the process data of the target training round and the reference value.
4. The method according to claim 1, wherein: The target processing is to skip the abnormal training round; the method further includes: The training samples of the target batch are put back into the training sample pool for resampling.
5. The method according to claim 1, wherein: The target processing is to adjust the hyperparameters in the abnormal training round; the hyperparameters include the learning rate; the target processing specifically includes: According to the target difference, the value of the learning rate is reduced, and the large language model is updated according to the adjusted learning rate.
6. The method according to claim 1, wherein: The target processing is to adjust the hyperparameters in the abnormal training round; the hyperparameters include the learning rate; the target processing specifically includes: The learning rate is reduced to a first preset value, and the large language model is updated according to the adjusted learning rate.
7. The method according to claim 6, further comprising: After the target processing, the training samples of the target batch are put back into the training sample pool for resampling.
8. The method according to claim 4 or 7, wherein: Putting the training samples of the target batch back into the training sample pool includes: Marking each training sample in the target batch as an abnormal training sample, and multiplying the sampling probability of each abnormal training sample by a preset first multiplier; the first multiplier is less than 1; The abnormal training samples are put back into the training sample pool so that they are sampled with their respective sampling probabilities in subsequent training rounds.
9. The method according to claim 1, wherein: The process data includes gradient values of various parameters; the method further includes: When the target difference does not exceed the first threshold, determining an average gradient value of each parameter according to the gradient value of each parameter in the N training rounds; Determine the gradient noise scale value according to the norm and covariance matrix determined by the average gradient values of the parameters; When the gradient noise scale value is greater than a preset second threshold, the target training round is determined as an abnormal training round.
10. The method according to claim 9, wherein: The average gradient value of each parameter is determined according to an exponential moving average.
11. The method according to claim 9, wherein: The target processing is to skip the abnormal training round; the method further includes: Increase the number of training samples used in each training round after the target training round.
12. The method according to claim 1 or 9, further comprising: When M consecutive training rounds are determined to be abnormal training rounds, the training process of the large language model is terminated.
13. A device for training a large language model, comprising: A process data determination unit is configured to determine process data of a target training round by inputting a target batch of training samples into the large language model, wherein the training samples include text data, and the process data include training loss values or gradient values of various parameters; An acquisition unit is configured to acquire a benchmark value obtained by statistically analyzing process data of N consecutive training rounds before a target training round; A first judgment unit is configured to determine the target training round as an abnormal training round when a target difference between the process data of the target training round and the reference value exceeds a preset first threshold; The abnormal processing unit is configured to perform target processing on the abnormal training round; the target processing includes skipping the abnormal training round, or adjusting the hyperparameters in the abnormal training round to reduce the impact of the abnormal training round.
14. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 12.
15. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 12 is implemented.