Large language model training method and device, equipment and storage medium

By calculating the cross entropy loss value in large language model training to screen high-quality corpus data, the problem of low training data quality is solved, and the performance and generation ability of the model are improved.

CN120258104APending Publication Date: 2025-07-04THE FIFTH AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242020.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

During the training process of existing large language models, due to the low quality of training data, the model performance improvement is limited.

Method used

By obtaining the corpus data training set, the loss value is calculated using the cross entropy loss function to determine whether the loss value is less than the preset value. If it is less than, the model parameters will not be updated. If it is greater than, the model parameters will be updated, and the corpus data will be reselected for training until the preset conditions are met, and the training data quality will be improved.

Benefits of technology

The performance of large language models is improved, and the quality of the corpus data is judged by calculating the loss value, thereby improving the quality of the training data, improving the generation ability of the model and prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258104A_ABST
    Figure CN120258104A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model training method and device, equipment and a storage medium, and the method comprises the steps: obtaining a to-be-trained corpus data training set, and selecting corpus data from the corpus data training set; predicting corpus data by using a large-scale language model to obtain a prediction result, calculating by using a cross entropy loss function based on the prediction result to obtain a loss value, storing the current loss value, calculating according to the stored current loss value to obtain a current preset loss value, judging whether the loss value is smaller than the current preset loss value or not, and if yes, judging whether the loss value is smaller than the current preset loss value or not. If the loss value is smaller than the current preset loss value, the model parameters are not updated, and if the loss value is larger than the current preset loss value, after the model parameters of the large-scale language model and the current preset loss value are updated according to the loss value, the corpus data are selected from the corpus data training set again to continue updating the model parameters of the large-scale language model until the trained large-scale language model is obtained. The loss value is calculated through the method, the corpus data quality is improved, and therefore the performance of a large language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and in particular to a method, device, equipment and storage medium for training large language models. Background Art

[0002] The pre-training of large language models (LLMs) is a complex and resource-intensive process, which involves using a large amount of data to train the model so that it can understand and generate natural language text. The pre-training stage is one of the core links in LLM training. In this stage, the model learns the statistical laws, semantic information and context relationships of the language by processing large-scale unlabeled text data. The pre-training dataset is usually extremely large, containing billions or even trillions of words, and this data comes from various sources such as the Internet, books, and academic papers. During the pre-training process, the model usually adopts the self-supervised learning method. For example, by predicting the masked or covered parts of the text to learn the associations between words and the structure of sentences. This method can make full use of a large amount of unlabeled data and improve the generalization ability of the model. However, the pre-training cost is very high and requires a large amount of computing resources and time. In addition, the inference latency of the model, the context length limit, and the instability of the evaluation method are also challenges that need to be faced during the pre-training process.

[0003] Currently, the main methods include: First, before starting the training, understand the required data types, data volume and quality requirements. Perform data cleaning to remove noise and outliers. Screen out high-quality data by constructing a special text scoring model. Second, use open-source and closed-source large models and pre-training data to let these large models generate high-quality data, and the pre-training data can be condensed into high-quality data to reduce the proportion of garbage data. Third, in Patent CN118171108A, the pre-training data of the model is scattered, and then divided into blocks and sorted according to the text data length, reducing the randomness of the dataset, improving the model training efficiency, and reducing the training cost. However, the above methods only simply perform data sorting and cannot improve the data quality, and have limited improvement on the performance of large models. Summary of the Invention

[0004] To solve the above technical problems, embodiments of the present invention provide a method, device, equipment and storage medium for training large language models, which solve the problem that the performance of the obtained large language model is low due to the low quality of training data during the training process of existing large language models.

[0005] The first aspect of the embodiments of the present invention provides a method for training a large language model, and the method includes: Obtain a corpus data training set to be trained, and select a piece of corpus data from the corpus data training set; Input the corpus data into a large language model for prediction to obtain a prediction result. Based on the prediction result, calculate the current loss value using the cross-entropy loss function, save the current loss value, and calculate the current preset loss value based on the saved current loss value; Determine whether the current loss value is less than the current preset loss value. If the current loss value is less than the current preset loss value, do not update the model parameters of the large language model; If the current loss value is greater than the current preset loss value, update the model parameters of the large language model according to the current loss value, re-select a new piece of corpus data from the corpus data training set, and continue to update the model parameters of the large language model based on the new corpus data until the updated model parameters meet the preset conditions to obtain a trained large language model.

[0006] In a possible implementation manner of the first aspect, calculating the current preset loss value based on the saved current loss value includes: Based on the saved multiple loss values, calculate the values of the exponential functions of each loss value to obtain multiple values; Perform a normalization operation on each value to obtain the normalized value corresponding to each value, and perform a descending order sorting on each normalized value to obtain a sorting result; Accumulate the normalized values ranked in the first N and the first N + 1 in the sorting result respectively to obtain a first accumulation result and a second accumulation result, where N is an integer greater than or equal to 1; Obtain the initial preset loss value according to the first accumulation result and the second accumulation result, and invert the initial preset loss value to obtain the current preset loss value.

[0007] In a possible implementation manner of the first aspect, accumulating the normalized values ranked in the first N and the first N + 1 in the sorting result respectively to obtain a first accumulation result and a second accumulation result includes: Select the normalized values ranked in the first N in the sorting result to obtain a first target sorting result, and select the normalized values ranked in the first N + 1 in the sorting result to obtain a second target sorting result; Determine whether there are normalized values greater than the current preset loss value in the first target sorting result and the second target sorting result. If so, remove the normalized values greater than the current preset loss value to obtain a new first target sorting result and a new second target sorting result; Add the normalized values in the new first target sorting result and the new second target sorting result respectively to obtain a first accumulation result and a second accumulation result.

[0008] In a possible implementation of the first aspect, obtaining an initial preset loss value based on the first accumulation result and the second accumulation result includes: Subtract the second accumulation result from the first accumulation result to obtain the initial preset loss value.

[0009] In a possible implementation of the first aspect, when re-selecting corpus data from the corpus data training set, it includes: Select un-trained corpus data from the corpus data training set to be trained; Or select corpus data that meets the preset conditions from the corpus data with a loss value less than the current preset loss value.

[0010] The second aspect of the embodiments of the present invention provides a large language model training device, including: An acquisition module, configured to acquire a corpus data training set to be trained and select a piece of corpus data from the corpus data training set; a loss value calculation module, configured to input the corpus data into a large language model for prediction to obtain a prediction result, calculate a current loss value based on the prediction result using a cross-entropy loss function, save the current loss value, and calculate a current preset loss value based on the saved current loss value; A first judgment module, configured to judge whether the current loss value is less than the current preset loss value. If the current loss value is less than the current preset loss value, the model parameters of the large language model are not updated; A second judgment module, configured to, if the current loss value is greater than the current preset loss value, update the model parameters of the large language model according to the current loss value, re-select a new piece of corpus data from the corpus data training set, and continue to update the model parameters of the large language model based on the new corpus data until the updated model parameters meet the preset conditions to obtain a trained large language model.

[0011] In a possible implementation of the second aspect, the judgment module includes a selection unit, a normalization unit, an accumulation unit, and a calculation unit, where The selection unit is configured to calculate the values of the exponential functions of each loss value based on the saved multiple loss values to obtain multiple values; the normalization unit is configured to perform a normalization operation on each value to obtain the normalized values corresponding to each value, and perform a descending order sorting on each normalized value to obtain a sorting result; The accumulation unit is configured to add the normalized values sorted in the first N and the first N + 1 in the sorting result to obtain a first accumulation result and a second accumulation result, where N is an integer greater than or equal to 1; The calculation unit is configured to obtain an initial preset loss value based on the first accumulation result and the second accumulation result, and invert the initial preset loss value to obtain the current preset loss value.

[0012] In a possible implementation of the second aspect, the accumulation unit includes a sorting subunit, a screening subunit, and a sorting result accumulation subunit. Among them, the sorting subunit is used to select the top N normalized values in the sorting result to obtain a first target sorting result, and select the top N+1 normalized values in the sorting result to obtain a second target sorting result; The screening subunit is used to determine whether the normalized values in the first target sorting result and the second target sorting result are greater than the current preset loss value. If so, remove the normalized values greater than the current preset loss value to obtain a new first target sorting result and a new second target sorting result; The sorting result accumulation subunit is used to add up the respective normalized values in the new first target sorting result and the new second target sorting result to obtain a first accumulation result and a second accumulation result.

[0013] A third aspect of the embodiments of the present invention provides a computer device, including: A memory for storing a computer program; A processor for implementing the large language model training method as in the first aspect when executing the computer program.

[0014] A fourth aspect of the embodiments of the present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the large language model training method as in the first aspect are implemented.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: The large language model training method provided by the embodiments of the present invention obtains a corpus data training set to be trained, and selects corpus data from the corpus data training set; uses the large language model to predict the corpus data to obtain a prediction result, and based on the prediction result, calculates a loss value using a cross-entropy loss function; determines whether the loss value is less than the current preset loss value. If the loss value is less than the current preset loss value, the model parameters of the large language model are not updated. If the loss value is greater than the current preset loss value, after updating the model parameters of the large language model and the current preset loss value, re-select corpus data from the corpus data training set, and continue to update the model parameters of the large language model based on the re-selected corpus data until the updated model parameters meet the preset conditions, and a trained large language model is obtained. Through the above method, by calculating the loss value of the corpus data to determine whether the corpus data can participate in the training, the quality of the training data is improved, and the performance of the large language model is improved. Description of the Drawings

[0016] Figure 1Flowchart of an embodiment of the large language model training method provided by the present invention; Figure 2 Flowchart of the large language model training method of an embodiment of the large language model training method provided by the present invention; Figure 3 Block diagram of another embodiment of the large language model training method provided by the present invention. Detailed implementation manners

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.

[0019] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0020] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0021] The term " / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0022] Therefore, a schematic flowchart of an embodiment of the large language model training method provided by the embodiments of the present invention is as follows Figure 1 , including steps S101 to S104, and the specific steps are as follows: S101. Obtain a corpus data training set to be trained, and select a piece of corpus data from the corpus data training set.

[0023] In this embodiment, pre-training data for the large language model is obtained from various sources such as the Internet, books, and academic papers to obtain a corpus data training set. The pre-training data set of the large language model is usually extremely large, containing billions or even trillions of words. The diversity and scale of the data are crucial for training a high-performance LLM. During training, a piece of corpus data is selected from the corpus data training set and input into the large language model.

[0024] S102. Input the corpus data into the large language model for prediction to obtain a prediction result. Based on the prediction result, calculate the current loss value using the cross-entropy loss function, save the current loss value, and calculate the current preset loss value based on the saved current loss value.

[0025] In this embodiment, as Figure 2 shown, the selected corpus data is input into the large language model for prediction to obtain a prediction result. Based on the prediction result, calculate the loss value using the cross-entropy loss function.

[0026] In the training process of most current large models, cross-entropy plays a crucial role. Cross-entropy is a commonly used loss function in natural language processing, especially in language models. It measures the difference between the probability distribution predicted by the model and the true distribution. In a large language model, the cross-entropy loss function is used to optimize the model parameters so that the model can better predict the probability of the next word or token. By minimizing the cross-entropy loss, the model learns to adjust its parameters to improve the prediction accuracy of the true data distribution.

[0027] The cross-entropy is defined for two discrete probability distributions P and Q, and the mathematical expression of the cross-entropy loss function is: In the formula, P represents the true probability distribution of the corpus data, Q represents the probability distribution of the prediction result of the large language model, and x represents the corpus data.

[0028] The large language model learns by predicting the next word in a sentence. During training, the model is made to minimize the difference between the prediction and the actual word. Therefore, cross-entropy acts as a loss function, guiding the model to adjust its parameters to produce higher-quality predictions. By minimizing the cross-entropy, the large language model becomes increasingly proficient at generating coherent, contextually relevant text that closely resembles human language patterns.

[0029] There are tens of billions or even trillions of corpora for model training in the general corpus data training set. During the model training and inference processes, after each corpus is calculated by the large language model to obtain a prediction result, the Loss value of the loss function obtained based on the prediction result is different for each corpus. Some corpora have a high Loss value, indicating a large difference between the prediction and the reality; some corpora have a low Loss value, indicating a low difference between the prediction and the reality.

[0030] However, it is found during the training of the large language model that the Loss value can be used as a criterion for judging the quality of the corpus. During the training process of the large language model, the large language model itself gradually discovers the rules of the corpus and gets familiar with the distribution of the corpus data. In this way, the model can calculate the loss function Loss value of a certain corpus to determine whether this corpus should participate in the training. If the loss function Loss is relatively low, it indicates that the quality of this corpus data is not high, and it may contain a lot of easily predictable data, the corpus contains a lot of formatted tokens or repeated tokens, or the change of the corpus tokens is relatively limited. Therefore, during training, the corpora with low Loss values can be skipped to improve the training performance of the large language model.

[0031] In some embodiments, "calculating the current preset loss value according to the saved current loss value" in step S102 includes but is not limited to the following steps: Based on the saved multiple loss values, calculate the values of the exponential functions of each loss value to obtain multiple values; Perform a normalization operation on each value to obtain the normalized value corresponding to each value, and perform a descending order sorting on each normalized value to obtain a sorting result; Accumulate the normalized values ranked in the first N and the first N + 1 in the sorting result respectively to obtain a first accumulation result and a second accumulation result, where N is an integer greater than or equal to 1; Obtain the initial preset loss value according to the first accumulation result and the second accumulation result, and invert the initial preset loss value to obtain the current preset loss value.

[0032] In this embodiment, during the training of the large language model, after obtaining the Loss value of each corpus, record this Loss value. In the code, the get_skip_value function can be used to obtain a numerical value skip_value, that is, the current preset loss value, based on the stored Loss value. If the Loss value of the currently trained corpus is lower than skip_value, then the model will not include this corpus in the backpropagation of the training this time, that is, this corpus data will not affect the change of the model training weights, which is equivalent to skipping this corpus and not including it in the current training consideration. If it is greater than skip_value, then the corresponding corpus of this item is included in the normal training.

[0033] It should be noted that the Loss value refers to the loss value obtained through the cross-entropy loss function.

[0034] Specifically, in the written pseudocode, the loss values of all the first t records in the record array are taken out, and the values of the exponential function with the natural constant e as the base are calculated. Then the obtained values are taken out and accumulated to obtain the accumulated value. Using the accumulated value, each value is normalized to obtain a new array, and the sum of the values in this array is 1. Then the normalized values in the new array are sorted in descending order to obtain the probs array, and the probs array stores the normalized values sorted in descending order.

[0035] Construct the cum_sum_probs function to accumulate the normalized values sorted in the top few positions in the probs array. For example: cum_sum_probs[3] = sort_probs[0] + sort_probs[1] + sort_probs[2] + sort_probs[3] Among them, cum_sum_probs[3] is the accumulated result of the normalized values sorted in the top 3 positions, and sort_probs[N] represents the normalized value ranked Nth.

[0036] After obtaining the accumulated result of the normalized values sorted in the top few positions, the initial preset loss value is obtained according to the accumulated result, and then the initial preset loss value is inverted to obtain the current preset loss value.

[0037] It should be noted that the parameter p value is the current preset loss value. By statistically analyzing the historical Loss pattern, a suitable skip_value value for the current model is selected. The parameter p is variable and generally takes a value between 0.5 and 0.8.

[0038] In some embodiments, the step "accumulate the normalized values sorted in the top N and top N + 1 positions in the sorting result respectively to obtain a first accumulated result and a second accumulated result" in step S103 includes but is not limited to the following steps: Select the normalized values sorted in the top N positions in the sorting result to obtain a first target sorting result, and select the normalized values sorted in the top N + 1 positions in the sorting result to obtain a second target sorting result; Judge whether there are any normalized values in the first target sorting result and the second target sorting result that are greater than the current preset loss value. If so, remove the normalized values greater than the current preset loss value to obtain a new first target sorting result and a new second target sorting result; Add the normalized values in the new first target sorting result and the new second target sorting result respectively to obtain a first accumulation result and a second accumulation result.

[0039] In this embodiment, select the top N normalized values in the sorting result for accumulation to obtain an accumulation result, that is, cum_sum_probs[N], and select the top N + 1 normalized values in the sorting result for accumulation to obtain an accumulation result, that is, cum_sum_probs[N+1]. When accumulating, use the input parameter p to split cum_sum_probs[N] and cum_sum_probs[N+1]. For example, when p = 0.8, discard the numbers greater than 0.8 in cum_sum_probs[N] and cum_sum_probs[N+1], and then obtain the final accumulation result.

[0040] In some embodiments, step S103 "obtaining an initial preset loss value according to the first accumulation result and the second accumulation result" includes but is not limited to the following steps: In this embodiment, after obtaining the final accumulation result of the top N normalized values and the final accumulation result of the top N + 1 normalized values, subtract the final accumulation result of the top N + 1 normalized values from the final accumulation result of the top N normalized values to obtain an initial preset loss value. For example, cum_sum_probs[1] - cum_sum_probs[2] can obtain a prob value, that is, the initial preset loss value. Then, reverse the initial preset loss value to a Loss value through prob_to_loss, that is, the current preset loss value, and finally return it.

[0041] S103. Determine whether the current loss value is less than the current preset loss value. If the current loss value is less than the current preset loss value, do not update the model parameters of the large language model.

[0042] In this embodiment, if the Loss value of the corpus data is less than skip_value, this corpus data is not included in the backpropagation of training, and the corpus with a lower Loss value is skipped.

[0043] S104. If the current loss value is greater than the current preset loss value, update the model parameters of the large language model according to the current loss value, select a new corpus data from the corpus data training set again, and continue to update the model parameters of the large language model based on the new corpus data until the updated model parameters meet the preset conditions to obtain a trained large language model.

[0044] In this embodiment, if the Loss value of the corpus data is greater than skip_value, then this corpus data is incorporated into the backpropagation of training to update the model parameters of the large language model. After obtaining the updated large language model, a new corpus data is selected again from the corpus data training set, and the model parameters of the large language model are continuously updated based on the new corpus data until the updated model parameters meet the preset conditions, obtaining the trained large language model.

[0045] In some embodiments, "when selecting corpus data again from the corpus data training set" in step S104 includes but is not limited to the following steps: Select untrained corpus data from the corpus data training set to be trained; Or select corpus data that meets the preset conditions from the corpus data with a loss value less than the current preset loss value.

[0046] In this embodiment, when selecting corpus data again from the corpus data training set, select untrained corpus data from the corpus data training set to be trained; Or select corpus data that meets the preset conditions from the corpus data with a loss value less than the current preset loss value, ensuring that a certain number of corpus data with a lower loss function Loss also participate in the training to ensure that the model does not forget some knowledge that has been learned.

[0047] Through comparative experimental analysis, in the same Llama LLM model based on the Transformer architecture, using the same corpus for pre-training, there is a relatively large increase in the evaluation set of the large model finally.

[0048] Table 1 shows the evaluation set scores of different large language model training methods The model used is the Llama architecture, and the size of the model is 1.5B (1.5 billion parameters). After the same corpus training and using the new algorithm, the scores in the evaluation set have all increased significantly. As shown in Table 1, Table 1 shows the evaluation set scores of each algorithm for training large language models.

[0049] It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1At least some of the steps may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed and completed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turns with at least some of the steps or stages in other steps or other steps.

[0050] In some embodiments, as Figure 3 shown, which shows a block diagram of a large language model training device 300 provided by an embodiment of the present application, including: an acquisition module 301, a loss value calculation module 302, a first judgment module 303, and a second judgment module 304. Among them, the acquisition module 301 is used to acquire a corpus data training set to be trained and select a piece of corpus data from the corpus data training set; The loss value calculation module 302 is used to input the corpus data into the large language model for prediction to obtain a prediction result, calculate the current loss value using a cross-entropy loss function based on the prediction result, save the current loss value, and calculate the current preset loss value according to the saved current loss value; The first judgment module 303 is used to judge whether the current loss value is less than the current preset loss value. If the current loss value is less than the current preset loss value, the model parameters of the large language model are not updated; The second judgment module 304 is used to, if the current loss value is greater than the current preset loss value, update the model parameters of the large language model according to the current loss value, re-select a new piece of corpus data from the corpus data training set, and continue to update the model parameters of the large language model based on the new corpus data until the updated model parameters meet the preset conditions to obtain a trained large language model.

[0051] In some embodiments, the judgment module includes a selection unit, a normalization unit, an accumulation unit, and a calculation unit. Among them, the selection unit is used to calculate the values of the exponential functions of each loss value based on the saved multiple loss values to obtain multiple values; the normalization unit is used to perform a normalization operation on each value to obtain the normalized value corresponding to each value, and perform a descending order sorting on each normalized value to obtain a sorting result; The accumulation unit is used to add the normalized values ranked in the first N and the first N + 1 in the sorting result to obtain a first accumulation result and a second accumulation result, where N is an integer greater than or equal to 1; The calculation unit is used to obtain an initial preset loss value according to the first accumulation result and the second accumulation result, and invert the initial preset loss value to obtain the current preset loss value.

[0052] In some embodiments, the accumulation unit includes a sorting subunit, a screening subunit, and a sorting result accumulation subunit, where The sorting subunit is used to select the top N normalized values in the sorting result to obtain the first target sorting result, and select the top N+1 normalized values in the sorting result to obtain the second target sorting result; The screening subunit is used to determine whether there are any normalized values in the first target sorting result and the second target sorting result that are greater than the current preset loss value. If so, remove the normalized values that are greater than the current preset loss value to obtain a new first target sorting result and a new second target sorting result; The sorting result accumulation subunit is used to add up the respective normalized values in the new first target sorting result and the new second target sorting result to obtain a first accumulation result and a second accumulation result.

[0053] The specific implementation manner of the apparatus for training a large language model is basically the same as the specific embodiments of the above-mentioned method for training a large language model, and will not be described in detail here.

[0054] In an embodiment of the present application, a computer device is provided. The computer device includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the above steps are implemented; for the computer device provided in this embodiment, its implementation principle and technical effects are similar to those of the above method embodiment, and will not be described in detail here.

[0055] In an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above steps are implemented; for the computer-readable storage medium provided in this embodiment, its implementation principle and technical effects are similar to those of the above method embodiment, and will not be described in detail here.

[0056] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0057] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. In particular, it is pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for training a large language model, characterized in that, Including: Obtain a training set of corpus data to be trained, and select a piece of corpus data from the training set of corpus data; Input the corpus data into a large language model for prediction to obtain a prediction result. Based on the prediction result, calculate the current loss value using the cross-entropy loss function, save the current loss value, and calculate the current preset loss value according to the saved current loss value; Judge whether the current loss value is less than the current preset loss value. If the current loss value is less than the current preset loss value, do not update the model parameters of the large language model; If the current loss value is greater than the current preset loss value, update the model parameters of the large language model according to the current loss value, re-select a new piece of corpus data from the training set of corpus data, and continue to update the model parameters of the large language model based on the new corpus data until the updated model parameters meet the preset conditions to obtain a trained large language model.

2. The large language model training method according to claim 1, wherein The calculating the current preset loss value according to the saved current loss value includes: Based on the saved multiple loss values, calculate the values of the exponential functions of each of the loss values to obtain multiple values; Perform a normalization operation on each of the values to obtain the normalized values corresponding to each of the values, and sort the normalized values in descending order to obtain a sorting result; Accumulate the normalized values ranked in the first N and the first N+1 in the sorting result respectively to obtain a first accumulation result and a second accumulation result, where N is an integer greater than or equal to 1; Obtain an initial preset loss value according to the first accumulation result and the second accumulation result, and invert the initial preset loss value to obtain the current preset loss value.

3. The large language model training method according to claim 2, wherein, The accumulating the normalized values ranked in the first N and the first N+1 in the sorting result respectively to obtain a first accumulation result and a second accumulation result includes: Select the normalized values ranked in the first N in the sorting result to obtain a first target sorting result, and select the normalized values ranked in the first N+1 in the sorting result to obtain a second target sorting result; Judge whether there are normalized values greater than the current preset loss value in the first target sorting result and the second target sorting result. If so, remove the normalized values greater than the current preset loss value to obtain a new first target sorting result and a new second target sorting result; Add up the normalized values in the new first target sorting result and the new second target sorting result respectively to obtain a first accumulation result and a second accumulation result.

4. The large language model training method according to claim 2, wherein, The obtaining the initial preset loss value according to the first accumulation result and the second accumulation result includes: Subtract the second accumulation result from the first accumulation result to obtain the initial preset loss value.

5. The large language model training method according to claim 1, wherein When re-selecting corpus data from the training set of corpus data, it includes: Select the corpus data that has not been trained from the training set of corpus data to be trained; Or select the corpus data that meets the preset conditions from the corpus data where the loss value is less than the current preset loss value.

6. A large language model training device, characterized in that, Including: An acquisition module, configured to acquire a corpus data training set to be trained, and select a piece of corpus data from the corpus data training set; A loss value calculation module, configured to input the corpus data into a large language model for prediction to obtain a prediction result, calculate a current loss value using a cross-entropy loss function based on the prediction result, save the current loss value, and calculate a current preset loss value based on the saved current loss value; A first judgment module, configured to judge whether the current loss value is less than the current preset loss value. If the current loss value is less than the current preset loss value, the model parameters of the large language model are not updated; A second judgment module, configured to, if the current loss value is greater than the current preset loss value, update the model parameters of the large language model according to the current loss value, re-select a new piece of corpus data from the corpus data training set, and continue to update the model parameters of the large language model based on the new corpus data until the updated model parameters meet the preset conditions to obtain a trained large language model.

7. The large language model training device according to claim 6, wherein The judgment module includes a selection unit, a normalization unit, an accumulation unit, and a calculation unit, where The selection unit is configured to calculate the values of the exponential functions of each of the loss values based on the saved multiple loss values to obtain multiple values; The normalization unit is configured to perform a normalization operation on each of the values to obtain the normalized values corresponding to each of the values, and perform a descending order sorting on each of the normalized values to obtain a sorting result; The accumulation unit is configured to add the normalized values ranked in the first N and the first N + 1 in the sorting result to obtain a first accumulation result and a second accumulation result, where N is an integer greater than or equal to 1; The calculation unit is configured to obtain an initial preset loss value according to the first accumulation result and the second accumulation result, and reverse the initial preset loss value to obtain the current preset loss value.

8. The large language model training device according to claim 7, wherein The accumulation unit includes a sorting subunit, a screening subunit, and a sorting result accumulation subunit, where The sorting subunit is configured to select the normalized values ranked in the first N in the sorting result to obtain a first target sorting result, and select the normalized values ranked in the first N + 1 in the sorting result to obtain a second target sorting result; the screening subunit is configured to judge whether there are any normalized values greater than the current preset loss value in the first target sorting result and the second target sorting result. If so, remove the normalized values greater than the current preset loss value to obtain a new first target sorting result and a new second target sorting result; The sorting result accumulation subunit is configured to add each of the normalized values in the new first target sorting result and the new second target sorting result respectively to obtain a first accumulation result and a second accumulation result.

9. A computer device, characterized in that, Including: A memory, configured to store a computer program; A processor for implementing the large language model training method according to any one of claims 1 to 5 when executing the computer program.

10. A storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by a processor, the large language model training method according to any one of claims 1 to 5 is implemented.