Language model training and data processing method and device, equipment and medium
By calculating and utilizing incremental loss values and relative loss values in the incremental fine-tuning process of the language model, iterative fine-tuning of the basic language model is solved, and the model's performance and adaptability in the target business is improved.
Patent Information
- Application Number
- CN202510179351.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
AI Technical Summary
There is a catastrophic forgetting problem during the incremental fine-tuning process. When learning new tasks, the model will forget the knowledge learned in the pre-training stage, resulting in the model's performance on the original tasks.
By obtaining the target business sample set, the incremental loss value and relative loss value are calculated for each business sample, and the pre-trained basic language model is iteratively fine-tuned based on these loss values to ensure that the model retains natural language comprehension capabilities while adapting to the target business.
It effectively reduces catastrophic forgetting, improves the performance and generalization capabilities of the language model in the target business, and enables the model to better adapt to multitasking learning.
Smart Images

Figure CN120012869A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing and deep learning technology, and in particular to a language model training and data processing method, device, equipment and medium. Background Art
[0002] In the field of Natural Language Processing (NLP), the rise of pre-trained language models (PLMs) marks a major leap forward in the field. These models successfully capture the deep features of language and contextual understanding capabilities through unsupervised pre-training on large-scale and diverse corpuses. The pre-training process enables the model to learn rich language representations, providing a solid foundation for subsequent specific tasks.
[0003] With the continuous development of technology, especially driven by advanced models such as ChatGPT, the NLP field has conducted in-depth research on the fine-tuning technique of pre-trained language models. Among them, supervised fine-tuning and incremental fine-tuning are two common fine-tuning methods. The supervised fine-tuning method aims to improve the comprehensive ability of the model on various tasks by further training the pre-trained model on a large-scale, multi-task dataset. Although this method has been successful to a certain extent, its high computational cost and time consumption have become a major obstacle in practical applications. In addition, supervised fine-tuning methods often find it difficult to achieve optimal performance in specific domains or tasks, limiting the flexibility and adaptability of the model.
[0004] In order to overcome the limitations of supervised fine-tuning, the incremental fine-tuning method came into being. This method introduces a small amount of data and computing resources to fine-tune the pre-trained model for targeted domains or task adaptability, thereby significantly improving the performance of the model on downstream tasks. The incremental fine-tuning method not only reduces the computing cost, but also improves the flexibility and adaptability of the model, enabling it to better adapt to various practical application scenarios.
[0005] However, the incremental fine-tuning method also faces a severe challenge: catastrophic forgetting. During the incremental fine-tuning process, the model tends to forget the knowledge learned in the pre-training phase when learning a new task. This phenomenon manifests itself in that while the performance of the model on the new task is improved, its performance on the original task is significantly reduced. The catastrophic forgetting problem severely limits the multi-task learning ability of the model, making it difficult for the model to maintain a good performance balance between multiple tasks.
[0006] Therefore, how to solve the problem of catastrophic forgetting during incremental fine-tuning has become a key technical problem that needs to be urgently solved in the current NLP field. Summary of the invention
[0007] The present application provides a language model training and data processing method, apparatus, device and medium for solving the problem of catastrophic forgetting in the existing process of incremental fine-tuning of language models.
[0008] In a first aspect, the present application provides a method for training a language model, the method comprising:
[0009] Acquire a target business sample set; wherein the target business sample set includes business samples and expected outputs of the business samples under the target business;
[0010] Based on the target business sample set, iteratively fine-tune the pre-trained basic language model; wherein the basic language model is a model that already has natural language understanding capabilities;
[0011] Among them, in any iterative fine-tuning process:
[0012] For any business sample, obtain a first prediction output and a first probability distribution corresponding to the business sample through the currently fine-tuned language model, and obtain a second prediction output and a second probability distribution corresponding to the business sample through the basic language model; determine an incremental loss value according to the first prediction output, the first probability distribution, and the expected output; and determine a relative loss value between the basic language model and the currently fine-tuned language model according to the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution;
[0013] Based on each of the incremental loss values and each of the relative loss values, the currently fine-tuned language model is fine-tuned to obtain a trained language model that supports the target business.
[0014] In a second aspect, the present application also provides a data processing method based on the above-mentioned model, the method comprising:
[0015] Obtaining the pending business data of the target business;
[0016] The processing result of the target business is obtained based on the business data to be processed through a pre-trained language model supporting the target business.
[0017] In a third aspect, the present application further provides a language model training device, the device comprising:
[0018] An acquisition module, used to acquire a target business sample set; wherein the target business sample set includes business samples and expected outputs of the business samples under the target business;
[0019] Training modules for:
[0020] Based on the target business sample set, iteratively fine-tune the pre-trained basic language model; wherein the basic language model is a model that already has natural language understanding capabilities;
[0021] Among them, in any iterative fine-tuning process:
[0022] For any business sample, obtain a first prediction output and a first probability distribution corresponding to the business sample through the currently fine-tuned language model, and obtain a second prediction output and a second probability distribution corresponding to the business sample through the basic language model; determine an incremental loss value according to the first prediction output, the first probability distribution, and the expected output; and determine a relative loss value between the basic language model and the currently fine-tuned language model according to the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution;
[0023] Based on each of the incremental loss values and each of the relative loss values, the currently fine-tuned language model is fine-tuned to obtain a trained language model that supports the target business.
[0024] In a fourth aspect, the present application further provides a data processing device based on the above-mentioned model, the device comprising:
[0025] An acquisition unit, used for acquiring the to-be-processed business data of the target business;
[0026] A processing unit is used to obtain a processing result of the target business based on the business data to be processed by using a pre-trained language model supporting the target business.
[0027] In a fifth aspect, the present application provides a computer device, comprising a processor, wherein the processor is used to implement the steps of the language model training method as described above when executing a computer program stored in a memory, or to implement the steps of the data processing method as described above.
[0028] In a sixth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the language model training method as described above, or implements the steps of the data processing method as described above.
[0029] The beneficial effects of this application are as follows:
[0030] By combining the incremental loss value and the relative loss value, the basic language model is iteratively fine-tuned so that the trained language model can effectively support specific target businesses. While adapting to the target business, this method retains the natural language understanding ability of the basic language model, improves the performance and generalization ability of the language model on the target business, and thus reduces catastrophic forgetting. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0032] Figure 1 A schematic diagram of a language model training process provided in an embodiment of the present application;
[0033] Figure 2 A schematic diagram of a flow chart for training a language model provided in an embodiment of the present application;
[0034] Figure 3 A schematic diagram of a data processing process provided by an embodiment of the present application;
[0035] Figure 4 A schematic diagram of the structure of a language model training device provided in an embodiment of the present application;
[0036] Figure 5 A schematic diagram of a data processing device structure provided in an embodiment of the present application;
[0037] Figure 6 It is a structural schematic diagram of a computer device provided in an optional embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0039] In order to avoid the catastrophic forgetting problem during the incremental fine-tuning of a language model and improve the robustness of the language model, the present application provides a language model training and data processing method, apparatus, device and medium.
[0040] Embodiment 1:
[0041] This application provides a method for training a language model. Figure 1 A schematic diagram of a language model training process provided in an embodiment of the present application, the process comprising:
[0042] S101: Acquire a target business sample set; wherein the target business sample set includes business samples and expected outputs of the business samples under the target business.
[0043] In the present application, the training method of the language model is applied to a computer device, which may be an intelligent terminal, such as a computer, a robot, etc., or a server, such as an application server, a business server, etc.
[0044] In order to improve the processing capability and quality of the language model under the target business, in this application, a target business sample set is collected in advance, so as to incrementally fine-tune the basic language model through the target business sample set, thereby obtaining a trained language model that supports the target business. Among them, the target business sample set is the basic data source for the entire training process. It contains business samples and the expected output of these business samples under the target business. Business samples are usually in the form of text data, and the expected output has different forms of expression depending on the target business.
[0045] For example, in the financial business scenario, business samples can be various financial news reports, such as "[Date], [Company Name] released its financial report for the previous quarter, with revenue increasing by [X]% year-on-year"; or they can be stock analysis reports, such as "Recently, [Stock Code] has experienced large stock price fluctuations due to the market environment and internal strategic adjustments of the company". For financial news samples, the corresponding expected output can be a judgment on the future development trend of the company in the financial news, such as "positive", "stable development", "risks exist", etc.; for stock analysis report samples, the corresponding expected output can be a prediction of the stock trend, such as "increase", "decrease", "remain flat". In the medical business scenario, business samples can be patients' medical records, including patient symptom descriptions, examination results and other information, such as "patient [name], age [X] years old, has had cough and fever symptoms in the past week, with a maximum temperature of [X]℃, and a blood test shows a white blood cell count of [X]". The expected output can be a preliminary diagnosis result given based on the medical record, such as "cold", "pneumonia", "bronchitis", etc. In the e-commerce business scenario, business samples can be user product reviews, such as "This phone has a stylish appearance and a good screen display, but the battery life is poor"; or they can be product search keywords. The expected output can be the sentiment classification of product reviews, such as "positive", "negative", or "neutral"; for search keywords, the expected output can be the recommended related product categories.
[0046] When obtaining business samples, data related to the target business can be collected from public data sets, internal business databases of enterprises, web crawlers, etc. For example, in financial business, data can be collected from financial news websites, official websites of stock exchanges, etc.; in medical business, medical record data can be obtained from the hospital's information system, but attention should be paid to data privacy protection and compliance. The collected data needs to be labeled to obtain the expected output. Among them, manual labeling can be used to organize professionals to label business samples according to the rules and requirements of the target business. For example, in text classification tasks, labelers classify business samples into corresponding categories according to their content. Some semi-automatic labeling tools can also be combined to improve labeling efficiency.
[0047] In a possible implementation, the collected raw data may be cleaned to remove interfering data.
[0048] S102: iteratively fine-tuning a pre-trained basic language model based on the target business sample set; wherein the basic language model is a model that already has natural language understanding capabilities;
[0049] Among them, in any iterative fine-tuning process:
[0050] For any business sample, obtain a first prediction output and a first probability distribution corresponding to the business sample through the currently fine-tuned language model, and obtain a second prediction output and a second probability distribution corresponding to the business sample through the basic language model; determine an incremental loss value according to the first prediction output, the first probability distribution, and the expected output; and determine a relative loss value between the basic language model and the currently fine-tuned language model according to the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution;
[0051] Based on each of the incremental loss values and each of the relative loss values, the currently fine-tuned language model is fine-tuned to obtain a trained language model that supports the target business.
[0052] After obtaining the target service sample set based on the above embodiment, the target service sample set can be used to iteratively fine-tune the pre-trained basic language model. The basic language model is a model that has been pre-trained based on massive text and has natural language understanding capabilities, such as BERT, GPT, Llama, etc.
[0053] It should be noted that when choosing a basic language model, it is necessary to consider factors such as model performance, applicable scenarios, and computing resources. For example, the GPT series of models perform well in text generation tasks, while the BERT series of models have good results in tasks such as text classification and question answering.
[0054] In this application, the parameters of the basic language model can be continuously optimized through multiple iterations to gradually adapt it to the target business. In each iteration, for each business sample in the target business sample set, the following operations need to be performed:
[0055] 1. Get prediction output and probability values
[0056] Input any business sample into the currently fine-tuned language model. The input business sample is processed by the currently fine-tuned language model to obtain the prediction output (recorded as the first prediction output) corresponding to the business sample and its corresponding probability value (recorded as the first probability distribution). Exemplarily, the business sample is input into the currently fine-tuned language model, and the currently fine-tuned language model will output a series of possible results, from which the result with the highest probability is selected as the first prediction output, and its corresponding first probability distribution is recorded.
[0057] At the same time, the business sample needs to be input into the basic language model. By processing the business sample through the basic language model, the prediction output (recorded as the second prediction output) corresponding to the business sample and its corresponding probability value (recorded as the second probability distribution) can be obtained. Subsequently, the current fine-tuned language model can be fine-tuned by comparing the performance difference between the current fine-tuned language model and the basic language model on the same business sample.
[0058] In a possible implementation, before inputting the business sample into the model (including the currently fine-tuned language model and the basic language model), the business sample may be segmented to convert it into an input format that the model can process. For example, a word segmenter is used to convert the text into a sequence of word vectors.
[0059] After obtaining the first predicted output corresponding to the business sample and its corresponding first probability distribution, as well as the second predicted output corresponding to the business sample and its corresponding second probability distribution based on the above embodiment, the incremental loss value can be calculated according to the first predicted output, the first probability distribution and the expected output. Among them, the incremental loss value reflects the gap between the prediction result of the current fine-tuned language model on the target business and the expected output. For example, the incremental loss value can be determined by loss functions such as cross entropy loss and mean square error loss.
[0060] In one example, determining the incremental loss value according to the first predicted output, the first probability distribution, and the expected output includes:
[0061] Determining a probability value that the first predicted output is consistent with the expected output according to the first predicted output, the first probability distribution, and the expected output;
[0062] Obtaining the natural logarithm of the probability value;
[0063] The inverse of the natural logarithm is determined as the incremental loss value.
[0064] When calculating the incremental loss value, we must first clarify the probability that the first predicted output is consistent with the expected output. Since the first probability distribution is the probability that the currently fine-tuned language model gives the first predicted output for the business sample, the probability value used to calculate the incremental loss value can be determined based on whether the first predicted output is the same as the expected output. Exemplarily, if the first predicted output and the expected output are exactly the same, then the probability value that the first predicted output is consistent with the expected output is equal to the probability corresponding to the first predicted output in the first probability distribution. For example, in a multi-class text sentiment analysis task, the expected output is "positive", the first predicted output is also "positive", and the first probability distribution is {'positive': 0.7, 'negative': 0.2, 'neutral': 0.1}, and the probability value of consistency is 0.7. If the first predicted output is different from the expected output, we can add the probabilities of all possible outputs except the first predicted output in the first probability distribution as the probability value that the first predicted output is consistent with the expected output. For example, the expected output is "negative", the first predicted output is "positive", and the first probability distribution is {'positive': 0.7, 'negative': 0.2, 'neutral': 0.1}, then the consistent probability value is 0.2+0.1=0.3.
[0065] After obtaining the probability value that the first predicted output is consistent with the expected output, it is necessary to take the natural logarithm. By taking the natural logarithm, the probability value can be mapped to a scale that is more conducive to calculation and analysis. Because the probability value is usually in the range of 0 to 1, directly using the probability value for loss calculation may cause problems such as numerical instability or gradient disappearance. Through the logarithmic function, the monotonically increasing characteristic of the logarithmic function can be used to better reflect the changes in the probability value in the subsequent optimization process.
[0066] Finally, the natural logarithm of the obtained probability value is taken inversely to obtain the incremental loss value. Since the logarithmic function is monotonically increasing, after taking the inverse, the larger the probability value, the smaller the loss value, which is in line with the goal of optimizing the model to improve the prediction accuracy.
[0067] For example, the incremental loss value is determined by the following formula:
[0068]
[0069] Among them, Loss 增量 represents the incremental loss value, W represents all tokens of the business sample, i represents the i-th token of the business sample, p(y i ″ =y i ) represents the first predicted output y of the i-th token i ″ and the expected output y i Consistent probability values.
[0070] In addition to the incremental loss value, it is also necessary to calculate the relative loss value between the basic language model and the currently fine-tuned language model. This loss value is determined based on the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution, and it reflects the degree of improvement of the fine-tuned language model relative to the basic language model in the target business. Among them, the calculation of the relative loss value can be based on a variety of methods, such as directly comparing the probability value difference of the two models on the same business sample, or comparing the average loss value difference of the two models on the same business sample set.
[0071] In a possible implementation, determining a relative loss value between the basic language model and the currently fine-tuned language model according to the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution includes:
[0072] Determine, according to the first predicted output, the first probability distribution, and the expected output, a first positive probability value that the first predicted output is consistent with the expected output, and a first negative probability value that the first predicted output is inconsistent with the expected output;
[0073] Determine, according to the second predicted output, the second probability distribution, and the expected output, a second positive probability value that the second predicted output is consistent with the expected output, and a second negative probability value that the second predicted output is inconsistent with the expected output;
[0074] Obtain a first positive product of the first positive probability value and the natural logarithm of the first positive probability ratio, and a first negative product of the first negative probability value and the natural logarithm of the first negative probability ratio; wherein the first positive probability ratio is the ratio of the first positive probability value to the second positive probability value, and the first negative probability ratio is the ratio of the first negative probability value to the second negative probability value;
[0075] Obtain a second positive product of the second positive probability value and the natural logarithm of the second positive probability ratio, and a second negative product of the second negative probability value and the natural logarithm of the second negative probability ratio; wherein the second positive probability ratio is the ratio of the second positive probability value to the first positive probability value, and the second negative probability ratio is the ratio of the second negative probability value to the first negative probability value;
[0076] determining a first relative sub-loss value of the currently fine-tuned language model according to a sum of the first positive product and the first negative product;
[0077] determining a second relative sub-loss value of the basic language model according to a sum of the second positive product and the second negative product;
[0078] The relative loss value is determined according to half of the sum of the second relative sub-loss value and the first relative sub-loss value.
[0079] In the process of determining the relative loss value, the first positive probability value and the first negative probability value must first be determined based on the first predicted output, the first probability distribution, and the expected output. Among them, the first positive probability value represents the probability that the first predicted output is consistent with the expected output, and the first negative probability value represents the probability that the first predicted output is inconsistent with the expected output. Assume that the first probability distribution is a dictionary, the key is the possible output, and the value is the corresponding probability. If the first predicted output is the same as the expected output, the first positive probability value is the probability corresponding to the first predicted output in the first probability distribution, and the first negative probability value is 1 minus the first positive probability value; if different, the first positive probability value is 1 minus the probability corresponding to the first predicted output in the first probability distribution, and the first negative probability value is the probability corresponding to the first predicted output in the first probability distribution.
[0080] Similarly, according to the second predicted output, the second probability distribution and the expected output, a second positive probability value and a second negative probability value are determined. It should be noted that the calculation logic of the second positive probability value and the second negative probability value is similar to the calculation logic of the first positive probability value and the first negative probability value, and will not be described in detail here.
[0081] Next, the ratio of the first positive probability value to the second positive probability value (recorded as the first positive probability ratio) and the ratio of the first negative probability value to the second negative probability value (recorded as the first negative probability ratio) are calculated. Then, the first positive probability value is multiplied by the natural logarithm of the first positive probability ratio to obtain a positive product (recorded as the first positive product), and the first negative probability value is multiplied by the natural logarithm of the first negative probability ratio to obtain a negative product (recorded as the first negative product).
[0082] Similarly, based on the similar acquisition method of the first positive product and the first negative product, the second positive product and the second negative product are acquired, which will not be described in detail here.
[0083] The first positive product and the first negative product are added to obtain a first relative sub-loss value of the currently fine-tuned language model, wherein the first relative sub-loss value reflects the degree of difference in output consistency between the currently fine-tuned language model and the basic language model.
[0084] Similarly, the second positive product and the second negative product are added to obtain a second relative sub-loss value of the basic language model. The second relative sub-loss value reflects the difference in output consistency between the basic language model and the currently fine-tuned language model.
[0085] Finally, the relative loss value is determined based on half of the sum of the second relative sub-loss value and the first relative sub-loss value. Through this relative loss value, the relative difference between the two models can be comprehensively considered to balance the impact on the base language model and the current fine-tuned language model during the fine-tuning process.
[0086] For example, the relative loss value can be expressed by the following formula:
[0087]
[0088] in, D2(yi″,yi′) and D1(y i ′ ,y i ″ ) is expressed in the same way, D1(y i ′ ,y i ″ ) represents the first relative sub-loss value, D2(y i ″ ,y i ′ ) represents the second relative sub-loss value, y i ,y i ′ ,y i ″ They represent the expected output, the first predicted output, and the second predicted output of the i-th token in the business sample, respectively. i ′ =y i ) represents the first positive probability value that the expected output of the i-th token in the business sample is consistent with the first predicted output, p(y i ′ ≠y i ) represents the first negative probability value that the expected output of the i-th token in the business sample is inconsistent with the first predicted output, p(y i ′ =y i )+p(y i ′ ≠y i )=1,p(y i ″ =y i) represents the second positive probability value that the expected output of the i-th token in the business sample is consistent with the second predicted output, p(y i ″ ≠y i ) represents the second negative probability value that the expected output of the i-th token in the business sample is inconsistent with the second predicted output, p(y i ″ =y i )+p(y i ″ ≠y i )=1.
[0089] Through the above embodiment, the calculation of the relative loss value between the basic language model and the currently fine-tuned language model can be completed. This relative loss value plays an important role in the iterative fine-tuning process of the language model, which helps to ensure that the fine-tuned model can adapt to the target business needs without deviating too much from the characteristics of the basic language model.
[0090] Through the above embodiment, several business samples can be obtained. For each business sample, the above operation is performed to obtain the relative loss value and incremental loss value of the business sample.
[0091] Finally, based on the incremental loss values and relative loss values of all business samples in the current iteration, the currently fine-tuned language model is further fine-tuned. For example, by updating the parameters of the currently fine-tuned language model, the prediction results of the language model on the target business are made more accurate.
[0092] In a possible implementation, each incremental loss value and each relative loss value may be weighted and summed to determine an overall loss value. Then, based on the overall loss value, the parameters of the currently fine-tuned language model are fine-tuned. This weighted summation method allows the user to adjust the importance of the incremental loss value and the relative loss value in the overall loss according to specific needs. Assuming that the weight of the incremental loss value is α and the weight of the relative loss value is β, the overall loss value can be expressed by the following formula:
[0093]
[0094] Among them, Loss represents the overall loss value, Loss 增量 Indicates the incremental loss value, Loss 相对 Represents the relative loss value, x represents the x-th business sample, and N represents the number of business samples.
[0095] Among them, the first weight value corresponding to the incremental loss value and the second weight value corresponding to the relative loss value can both be 1.
[0096] When the preset convergence conditions are met, the basic language model training is completed and an optimized language model is obtained. Among them, the preset convergence conditions can be met for the sum of the overall loss values determined based on each business sample to be less than a pre-configured loss threshold, or the determined overall loss value has been in a downward trend and tends to be flat, or the number of iterations for training the basic language model reaches the set maximum number of iterations, etc. It can be flexibly set in the specific implementation and is not specifically limited here.
[0097] As a possible implementation method, when training the basic language model, the business samples can be divided into training samples and test samples. The basic language model is first trained based on the training samples, and then the reliability of the trained language model is verified based on the test samples.
[0098] Figure 2 A flow chart of a language model training process provided for an embodiment of the present application. The language model training method of the present application is intended to enable a basic language model that already has natural language understanding capabilities to better adapt to specific target businesses. By iteratively fine-tuning using a target business sample set, while taking into account the incremental loss value of the current fine-tuned language model, and the relative loss value between the current fine-tuned language model and the basic language model, the trained language model not only meets the target business needs, but also retains the general natural language understanding capabilities of the basic language model.
[0099] The beneficial effects of this application are as follows:
[0100] By combining the incremental loss value and the relative loss value, the basic language model is iteratively fine-tuned so that the trained language model can effectively support specific target businesses. While adapting to the target business, this method retains the natural language understanding ability of the basic language model, improves the performance and generalization ability of the language model on the target business, and thus reduces catastrophic forgetting.
[0101] Embodiment 2:
[0102] The present application provides a data processing method based on the language model trained by the above embodiment. Figure 3 A schematic diagram of a data processing process provided in an embodiment of the present application, the process includes:
[0103] S301: Obtain the to-be-processed business data of the target business.
[0104] S302: Obtaining a processing result of the target business based on the business data to be processed by using a pre-trained language model supporting the target business.
[0105] The data processing method provided in the present application is applied to a computer device, which may be an intelligent device or a server. The computer device for data processing in the present application may be the same as or different from the computer device for language model training.
[0106] In a possible implementation, incremental training of the basic language model is generally performed offline. After the basic language model training is completed and a language model supporting the target service is obtained, the language model can be saved in the computer device for data processing.
[0107] Computer equipment can continuously collect the business data to be processed under the target business. The form and source of the business data to be processed will vary according to different target businesses.
[0108] After obtaining the business data to be processed, it is input into a pre-trained language model that supports the target business, thereby obtaining the processing result of the target business.
[0109] In one example, before the business data to be processed is input into the language model, some preprocessing operations are usually required to ensure that the format and quality of the data meet the requirements of the model. Common preprocessing steps include word segmentation, removal of stop words, and conversion to an input format acceptable to the model.
[0110] It should be noted that the specific training process of the language model has been described in the above embodiment, and the repeated parts will not be repeated here.
[0111] Embodiment 3:
[0112] Based on the same inventive concept, the present application also provides a language model training device, Figure 4 A schematic diagram of a language model training device provided in an embodiment of the present application, the device comprising:
[0113] The acquisition module 41 is used to acquire a target business sample set; wherein the target business sample set includes business samples and expected outputs of the business samples under the target business;
[0114] The training module 42 is used to:
[0115] Based on the target business sample set, iteratively fine-tune the pre-trained basic language model; wherein the basic language model is a model that already has natural language understanding capabilities;
[0116] Among them, in any iterative fine-tuning process:
[0117] For any business sample, obtain a first prediction output and a first probability distribution corresponding to the business sample through the currently fine-tuned language model, and obtain a second prediction output and a second probability distribution corresponding to the business sample through the basic language model; determine an incremental loss value according to the first prediction output, the first probability distribution, and the expected output; and determine a relative loss value between the basic language model and the currently fine-tuned language model according to the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution;
[0118] Based on each of the incremental loss values and each of the relative loss values, the currently fine-tuned language model is fine-tuned to obtain a trained language model that supports the target business.
[0119] The language model training device in this embodiment is presented in the form of a functional module, where the module refers to an application specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0120] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0121] Embodiment 4:
[0122] The present application also provides a data processing device based on the language model trained in the above embodiment 1, Figure 5 A schematic diagram of a data processing device structure provided in an embodiment of the present application, the device comprising:
[0123] An acquisition unit 51 is used to acquire the to-be-processed business data of the target business;
[0124] The processing unit 52 is used to obtain a processing result of the target business based on the business data to be processed by using a pre-trained language model supporting the target business.
[0125] The data processing device in this embodiment is presented in the form of a functional module, where the module refers to an application specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0126] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0127] Embodiment 5:
[0128] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present application, such as Figure 6 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0129] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0130] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0131] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0132] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0133] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.
[0134] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0135] Embodiment 6:
[0136] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor implements the following steps when executing:
[0137] Acquire a target business sample set; wherein the target business sample set includes business samples and expected outputs of the business samples under the target business;
[0138] Based on the target business sample set, iteratively fine-tune the pre-trained basic language model; wherein the basic language model is a model that already has natural language understanding capabilities;
[0139] Among them, in any iterative fine-tuning process:
[0140] For any business sample, obtain a first prediction output and a first probability distribution corresponding to the business sample through the currently fine-tuned language model, and obtain a second prediction output and a second probability distribution corresponding to the business sample through the basic language model; determine an incremental loss value according to the first prediction output, the first probability distribution, and the expected output; and determine a relative loss value between the basic language model and the currently fine-tuned language model according to the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution;
[0141] Based on each of the incremental loss values and each of the relative loss values, the currently fine-tuned language model is fine-tuned to obtain a trained language model that supports the target business.
[0142] Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to the training method of the language model, the implementation of the above-mentioned computer-readable storage medium can refer to Example 1 of the method, and the repeated parts will not be repeated.
[0143] Embodiment 7:
[0144] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor implements the following steps when executing:
[0145] Obtaining the pending business data of the target business;
[0146] The processing result of the target business is obtained based on the business data to be processed through a pre-trained language model supporting the target business.
[0147] Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to that of the data processing method, the implementation of the above-mentioned computer-readable storage medium can refer to Example 2 of the method, and the repeated parts will not be repeated.
[0148] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A method for training a language model, characterized in that: The method comprises: Acquire a target business sample set; wherein the target business sample set includes business samples and expected outputs of the business samples under the target business; Based on the target business sample set, iteratively fine-tune the pre-trained basic language model; wherein the basic language model is a model that already has natural language understanding capabilities; Among them, in any iterative fine-tuning process: For any business sample, obtain a first prediction output and a first probability distribution corresponding to the business sample through the currently fine-tuned language model, and obtain a second prediction output and a second probability distribution corresponding to the business sample through the basic language model; determine an incremental loss value according to the first prediction output, the first probability distribution, and the expected output; and determine a relative loss value between the basic language model and the currently fine-tuned language model according to the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution; Based on each of the incremental loss values and each of the relative loss values, the currently fine-tuned language model is fine-tuned to obtain a trained language model that supports the target business.
2. The method according to claim 1, characterized in that The determining the incremental loss value according to the first predicted output, the first probability distribution, and the expected output includes: Determining a probability value that the first predicted output is consistent with the expected output according to the first predicted output, the first probability distribution, and the expected output; Obtaining the natural logarithm of the probability value; The inverse of the natural logarithm is determined as the incremental loss value.
3. The method according to claim 1, characterized in that The determining, according to the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution, a relative loss value between the basic language model and the currently fine-tuned language model comprises: Determine, according to the first predicted output, the first probability distribution, and the expected output, a first positive probability value that the first predicted output is consistent with the expected output, and a first negative probability value that the first predicted output is inconsistent with the expected output; Determine, according to the second predicted output, the second probability distribution, and the expected output, a second positive probability value that the second predicted output is consistent with the expected output, and a second negative probability value that the second predicted output is inconsistent with the expected output; Obtain a first positive product of the first positive probability value and the natural logarithm of the first positive probability ratio, and a first negative product of the first negative probability value and the natural logarithm of the first negative probability ratio; wherein the first positive probability ratio is the ratio of the first positive probability value to the second positive probability value, and the first negative probability ratio is the ratio of the first negative probability value to the second negative probability value; Obtain a second positive product of the second positive probability value and the natural logarithm of the second positive probability ratio, and a second negative product of the second negative probability value and the natural logarithm of the second negative probability ratio; wherein the second positive probability ratio is the ratio of the second positive probability value to the first positive probability value, and the second negative probability ratio is the ratio of the second negative probability value to the first negative probability value; determining a first relative sub-loss value of the currently fine-tuned language model according to a sum of the first positive product and the first negative product; determining a second relative sub-loss value of the basic language model according to a sum of the second positive product and the second negative product; The relative loss value is determined according to half of the sum of the second relative sub-loss value and the first relative sub-loss value.
4. The method according to claim 1, characterized in that The step of fine-tuning the currently fine-tuned language model based on each of the incremental loss values and each of the relative loss values includes: Performing a weighted summation on the incremental loss values and the relative loss values to determine an overall loss value; Based on the overall loss value, parameters of the currently fine-tuned language model are fine-tuned.
5. The method according to claim 4, characterized in that The first weight value corresponding to the incremental loss value and the second weight value corresponding to the relative loss value are both 1.
6. A data processing method based on a language model trained by the method according to any one of claims 1 to 5, characterized in that: The method comprises: Obtaining the pending business data of the target business; The processing result of the target business is obtained based on the business data to be processed through a pre-trained language model supporting the target business.
7. A language model training device, characterized in that: The device comprises: An acquisition module, used to acquire a target business sample set; wherein the target business sample set includes business samples and expected outputs of the business samples under the target business; Training modules for: Based on the target business sample set, iteratively fine-tune the pre-trained basic language model; wherein the basic language model is a model that already has natural language understanding capabilities; Among them, in any iterative fine-tuning process: For any business sample, obtain a first prediction output and a first probability distribution corresponding to the business sample through the currently fine-tuned language model, and obtain a second prediction output and a second probability distribution corresponding to the business sample through the basic language model; determine an incremental loss value according to the first prediction output, the first probability distribution, and the expected output; and determine a relative loss value between the basic language model and the currently fine-tuned language model according to the first prediction output, the first probability distribution, the second prediction output, and the second probability distribution; Based on each of the incremental loss values and each of the relative loss values, the currently fine-tuned language model is fine-tuned to obtain a trained language model that supports the target business.
8. A data processing device based on a language model trained by the method according to any one of claims 1 to 5, characterized in that: The device comprises: An acquisition unit, used for acquiring the to-be-processed business data of the target business; A processing unit is used to obtain a processing result of the target business based on the business data to be processed by using a pre-trained language model supporting the target business.
9. A computer device, characterized in that: The computer device includes a processor, which is used to implement the steps of the language model training method as described in any one of claims 1 to 5 above when executing a computer program stored in a memory, or to implement the steps of the data processing method as described in claim 6 above.
10. A computer-readable storage medium, characterized in that: It stores a computer program that can be executed by a computer device. When the program is run on the computer device, the computer device executes the steps of the language model training method as described in any one of claims 1 to 5 above, or implements the steps of the data processing method as described in claim 6 above.