Model training method and device, equipment, storage medium and program product
By combining the training mechanism of self-supervised learning and reinforcement learning, using low-cost, label-free data screening and optimization of inference results, the problems of high training cost and low efficiency in the existing technology are solved, and the efficient inference ability of large language models is improved.
Patent Information
- Application Number
- CN202510337577.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The method of training large language models (LLM) in the prior art relies on a large amount of manual annotation data, resulting in high training costs and limited inference capabilities, self-supervised learning has limited performance in complex tasks, while reinforcement learning has low training efficiency and high computing resources.
Combined with the training mechanism of self-supervised learning and reinforcement learning, we obtain label-free data at low cost, screen reasonable inference results and conduct model training, and use the self-supervised learning stage to improve model reasoning capabilities. In the reinforcement learning stage, we optimize the model through high-quality feedback information to achieve rapid and steady improvement.
It reduces the cost of model training, improves the reasoning ability of LLM, realizes excellent reasoning ability in various tasks, and improves training efficiency and resource utilization efficiency.
Smart Images

Figure CN120258078A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Artificial Intelligence (AI) technology, and in particular, to a model training method, apparatus, device, storage medium, and program product. Background Art
[0002] A Large Language Model (LLM) is an artificial intelligence model based on deep learning technology. It is dedicated to understanding and generating human language and can complete various tasks such as dialogue, question answering, translation, and writing. It is one of the core technologies in the field of artificial intelligence.
[0003] In recent years, with the rapid development of LLM and deep learning technology, how to train LLM more efficiently and intelligently has become a hot research topic in the industry. Traditional model training methods (such as Supervised Learning) need to rely on a large amount of manually labeled training data to train LLM. The acquisition cost of training data is high, and in many cases, it is difficult to obtain sufficient training data, resulting in limited reasoning ability of the trained LLM. Summary of the Invention
[0004] Embodiments of this application provide a model training method, apparatus, device, storage medium, and program product, which can use a large amount of training data obtained at low cost to train LLM, thereby reducing the model training cost and improving the reasoning ability of the trained LLM.
[0005] In a first aspect of this application, a model training method is provided. The method includes:
[0006] According to first input data, determine multiple first inference results through a first language model. The first inference results include a first inference path of the first language model for the first input data and a first inference conclusion generated by the first language model based on the first input data;
[0007] Determine reasonable inference results among the multiple first inference results according to a preset discrimination rule;
[0008] Train the first language model based on the reasonable inference results to obtain a second language model;
[0009] According to second input data, determine multiple second inference results through the second language model. The second inference results include a second inference path of the second language model for the second input data and a second inference conclusion generated by the second language model based on the second input data;
[0010] Determine the evaluation results corresponding to each of the multiple second inference results;
[0011] Determine a reference inference result among the multiple second inference results according to the evaluation results corresponding to each of the multiple second inference results;
[0012] Train the second language model based on the reference inference result to obtain a target language model.
[0013] The second aspect of the present application provides a model training device, and the device includes:
[0014] A first inference module, configured to determine multiple first inference results through a first language model according to first input data, where the first inference results include a first inference path of the first language model for the first input data and a first inference conclusion generated by the first language model based on the first input data;
[0015] A first screening module, configured to determine a reasonable inference result among the multiple first inference results according to a preset discrimination rule;
[0016] A first training module, configured to train the first language model based on the reasonable inference result to obtain a second language model;
[0017] A second inference module, configured to determine multiple second inference results through the second language model according to second input data, where the second inference results include a second inference path of the second language model for the second input data and a second inference conclusion generated by the second language model based on the second input data;
[0018] An evaluation module, configured to determine the evaluation results corresponding to each of the multiple second inference results;
[0019] A second screening module, configured to determine a reference inference result among the multiple second inference results according to the evaluation results corresponding to each of the multiple second inference results;
[0020] A second training module, configured to train the second language model based on the reference inference result to obtain a target language model.
[0021] The third aspect of the present application provides a computer device, and the device includes a processor and a memory:
[0022] The memory is used to store a computer program;
[0023] The processor is configured to execute the steps of the model training method as described in the first aspect above according to the computer program.
[0024] The fourth aspect of this application provides a computer-readable storage medium for storing a computer program for executing the steps of the model training method described in the first aspect above.
[0025] The fifth aspect of this application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the steps of the model training method described in the first aspect above.
[0026] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0027] The embodiment of the present application provides a model training method, which innovatively proposes a mechanism for training an LLM by combining self-supervised learning and reinforcement learning. In the self-supervised learning stage, a large amount of unlabeled first input data can be obtained at low cost; then, through a first language model, a variety of first inference results are determined according to the first input data; and according to a preset discrimination rule, relatively accurate reasonable inference results are selected from the variety of first inference results; furthermore, based on the selected reasonable inference results, the first language model is trained to enable the first language model to initially learn how to reason reasonably based on the input data and improve the inference ability of the first language model. After the above self-supervised learning stage, a second language model with relatively good inference ability will be obtained, and this second language model will be used as the training object in the reinforcement learning stage. Thus, by enabling the training object in the reinforcement learning stage to initially have relatively good performance, it helps to improve the training efficiency in the reinforcement learning stage. In the reinforcement learning stage, a large amount of unlabeled second input data can still be obtained at low cost; then, through the second language model, a variety of second inference results are determined according to the second input data. Since the second language model is trained by self-supervised learning, generally, the quality of the second inference results generated by this second language model is relatively high; then, evaluation results corresponding to the variety of second inference results are determined to give corresponding feedback information for the reinforcement learning stage through the evaluation results corresponding to the variety of second inference results; furthermore, according to the feedback information represented by the evaluation results corresponding to the variety of second inference results, reference inference results participating in the reinforcement learning training are determined from the variety of second inference results, that is, from the relatively high-quality second inference results generated by the second language model trained in the self-supervised learning stage, reference inference results for guiding the model training direction in the reinforcement learning stage are further determined; finally, based on the reference inference results, the second language model is trained to enable the second language model to further learn how to reason more accurately and generate better-quality inference results from the reference inference results, and deeply optimize the inference ability of this second language model to obtain the target language model. From the above self-supervised learning stage and reinforcement learning stage, it can be seen that the present application can use a large amount of unlabeled data obtained at low cost to train the LLM, reduce the model training cost, and jointly train the LLM in two learning stages, which can enable the inference ability of the LLM to be improved rapidly and steadily. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0029] Figure 1It is the architecture diagram of the computer system provided by the embodiment of the present application;
[0030] Figure 2 It is the schematic diagram of the application scenario of the model training method provided by the embodiment of the present application;
[0031] Figure 3 It is the schematic flow diagram of the model training method provided by the embodiment of the present application;
[0032] Figure 4 It is the schematic diagram for determining the evaluation scores corresponding to various second inference results provided by the embodiment of the present application;
[0033] Figure 5 It is the schematic diagram for training the evaluation model provided by the embodiment of the present application;
[0034] Figure 6 It is the schematic diagram for determining the reference inference result provided by the embodiment of the present application;
[0035] Figure 7 It is the schematic diagram for training the second language model provided by the embodiment of the present application;
[0036] Figure 8 It is the overall architecture schematic diagram of the model training method provided by the embodiment of the present application
[0037] Figure 9 It is the schematic diagram of the news low-quality residue rate provided by the embodiment of the present application;
[0038] Figure 10 It is the structural schematic diagram of the model training device provided by the embodiment of the present application;
[0039] Figure 11 It is the structural schematic diagram of the terminal device provided by the embodiment of the present application;
[0040] Figure 12 It is the structural schematic diagram of the server provided by the embodiment of the present application. Detailed implementation manners
[0041] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0042] In the description, claims and the above-mentioned drawings of the present application, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0043] It should also be noted that before and during the process of collecting relevant data of the user, the present application can display a prompt interface, a pop-up window or output voice prompt information. The prompt interface, pop-up window or voice prompt information is used to prompt the user that their relevant data is being collected currently, so that the present application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the confirmation operation of the user on the prompt interface or the pop-up window. Otherwise (that is, when the confirmation operation of the user on the prompt interface or the pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. In other words, all user data collected by the present application is collected with the consent and authorization of the user, and the collection, use and processing of the relevant user data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0044] In the related art, an LLM can be trained through supervised learning, or self-supervised learning, or reinforcement learning. Currently, most solutions use supervised learning to train an LLM. When training an LLM based on supervised learning, it is necessary to rely on a large amount of manually labeled training data, and the acquisition cost of the training data is high, and the performance of the trained LLM is also restricted by the high acquisition cost of the training data.
[0045] Self-supervised learning does not need to rely on manually labeled training data. It learns the internal structure of the input data and the potential relationship between data by predicting some parts of the input data (such as predicting the masked words in a sentence, etc.). Although self-supervised learning can achieve excellent performance in many tasks without relying on manually labeled training data, its performance in complex tasks is still limited. Self-supervised learning usually cannot guide the LLM to perform in-depth optimization of the task objective. When dealing with complex reasoning tasks, relying only on self-supervised learning for model training cannot make the reasoning ability and accuracy of the LLM meet the required standards.
[0046] Reinforcement learning enables the LLM to interact with the environment to learn how to select optimal actions in specific states, thereby maximizing cumulative rewards. Different from the previous two training methods, reinforcement learning uses a trial-and-error mechanism and reward signals to guide model optimization. Reinforcement learning relies on the trial-and-error mechanism for exploration. Each feedback adjustment requires a large amount of computational resources and time, and most reinforcement learning algorithms need to go through a large number of repeated trainings to obtain effective strategies, resulting in low training efficiency and high consumption of computational resources for reinforcement learning. Moreover, the feedback mechanism of reinforcement learning is relatively single, relying only on the processing results of tasks for rewards and ignoring the detailed optimization in the parsing process, which also leads to suboptimal processing effects in some tasks.
[0047] To solve the problems existing in the above related technologies, an embodiment of the present application provides a model training method, which can train an LLM by combining self-supervised learning and reinforcement learning. In the self-supervised learning stage, a large amount of unlabeled first input data can be obtained at low cost; then, through a first language model, various first inference results are determined according to the first input data; and according to a preset discrimination rule, relatively accurate reasonable inference results are screened out from the various first inference results; furthermore, based on the screened reasonable inference results, the first language model is trained so that the first language model initially learns how to reasonably infer based on the input data, and the inference ability of the first language model is improved. After the above self-supervised learning stage, a second language model with better inference ability will be obtained, and this second language model will be used as the training object in the reinforcement learning stage. Thus, by making the training object in the reinforcement learning stage initially have better performance, the time required for the model to explore effective strategies in the reinforcement learning stage is shortened, thereby helping to improve the training efficiency in the reinforcement learning stage and reducing the consumption of computing resources. In the reinforcement learning stage, a large amount of unlabeled second input data can still be obtained at low cost; then, through the second language model, various second inference results are determined according to the second input data. Since the second language model is trained through self-supervised learning, the second inference results generated by this second language model generally have high quality; then, the evaluation results corresponding to the various second inference results are determined to give corresponding feedback information for the reinforcement learning stage through the evaluation results corresponding to the various second inference results; furthermore, according to the feedback information represented by the evaluation results corresponding to the various second inference results, reference inference results for participating in the reinforcement learning training are determined from the various second inference results, that is, from the relatively high-quality second inference results generated by the second language model trained through the self-supervised learning stage, reference inference results for guiding the model training direction in the reinforcement learning stage are further determined; finally, based on the reference inference results, the second language model is trained so that the second language model further learns how to make more accurate inferences and generate higher-quality inference results from the reference inference results, deeply optimizing the inference ability of this second language model to obtain a target language model, so that the trained target language model has better inference ability in various tasks.
[0048] To facilitate the understanding of the model training method provided by the embodiment of the present application, the computer system for implementing this model training method will be exemplarily introduced below.
[0049] See Figure 1 , Figure 1 which is the architecture diagram of the computer system provided by the embodiment of the present application. Figure 1The computer system 100 therein is a system architecture for implementing the model training method. The computer system 100 includes a terminal device 120, a database 130, and a server 140.
[0050] An application for triggering the training of the LLM is installed and running in the terminal device 120. The user can trigger the operation of training the LLM through this application in the terminal device 120. In response to the trigger operation, the terminal device 120 can send a request to the server 140. After receiving the request, the server 140 can execute the model training method in the embodiments of the present application to train the LLM and obtain the target language model.
[0051] The terminal device 120 can be connected to the server 140 through a wireless network or a wired network.
[0052] The server 140 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud computing services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0053] The server 140 can access the database 130 through the network, or the database 130 can also be integrated inside the server 140. The database 130 is used to store the first input data and the second input data.
[0054] Exemplarily, the server 140 includes a processor 144 and a memory 142. Among them, the memory 142 is used to store computer programs, such as computer programs for executing the model training method in the embodiments of the present application. The processor 144 is used to read the computer program from the memory 142 to execute the model training method in the embodiments of the present application. Specifically, when the processor 144 executes the model training method, it can obtain the first input data and the second input data from the database 130; after obtaining the first input data, it can determine multiple first inference results through the first language model according to the first input data; then, it can determine the reasonable inference result among the multiple first inference results according to the preset discrimination rule; based on the reasonable inference result, training the first language model can obtain the second language model; further, according to the second input data, it can determine multiple second inference results through the second language model; furthermore, it can determine the evaluation results corresponding to the multiple second inference results respectively. According to the evaluation results corresponding to the multiple second inference results respectively, it can determine the reference inference result among the multiple second inference results; finally, based on the reference inference result, training the second language model can obtain the target language model.
[0055] Optionally, the server 140 undertakes the main computing work, and the terminal device 120 undertakes the secondary computing work; or, the server 140 undertakes the secondary computing work, and the terminal device 120 undertakes the main computing work; or, the server 140 and the terminal device 120 adopt a distributed computing architecture for collaborative computing.
[0056] The embodiments of the present application do not limit the form of the applications installed on the terminal device 120, including but not limited to the client (Application, App), applet, etc. installed in the terminal device 120, and can also be in the form of a web page. The terminal device 120 can generally refer to one of multiple terminal devices, and only the terminal device 120 is used as an example in this embodiment. The device types of the terminal device 120 include but are not limited to at least one of a smart phone, a tablet computer, a wearable device, a personal computer (Personal Computer, PC), a laptop computer, and a desktop computer.
[0057] Those skilled in the art can know that the numbers of the above terminal devices 120 and the database 130 can be more or less. For example, the above terminal device 120 can be only one, or the above terminal device 120 can be multiple or a larger number. Correspondingly, the database 130 can also be only one, or the above database 130 can be multiple or a larger number. The embodiments of the present application do not limit the numbers and device types of the terminal device 120 and the database 130.
[0058] See Figure 2 , Figure 2 which is a schematic diagram of the application scenario of the model training method provided by the embodiments of the present application. This method is executed by a computer device, and this computer device can be, for example, Figure 1 the server 140 shown. The steps of the model training method executed by the server 140 are as follows:
[0059] The server 140 can obtain the first input data 201 and the second input data 202 from the database 130.
[0060] After the server 140 obtains the first input data 201, the first input data 201 can be input into the first language model 203 for processing to obtain multiple first inference results (such as Figure 2 the first inference result 1 204 shown,..., the first inference result N 205), and the first inference result includes the first inference path of the first language model 203 for the first input data 201 and the first inference conclusion generated by the first language model 203 based on the first input data 201.
[0061] After that, according to the preset discrimination rules, a reasonable inference result 206 can be determined from multiple first inference results. Based on the reasonable inference result 206, the first language model 203 can be trained to obtain the second language model 207.
[0062] Then, the second input data 202 can be input into the second language model 207 for processing to obtain multiple second inference results (such as Figure 2 the second inference result 1 208 shown,..., the second inference result M 209). The second inference result includes the second inference path of the second language model 207 for the second input data 202 and the second inference conclusion generated by the second language model 207 based on the second input data 202.
[0063] Furthermore, the evaluation results corresponding to the multiple second inference results can be determined (such as Figure 2 the evaluation result 1 210 corresponding to the second inference result 1 208 shown,..., the evaluation result M 211 corresponding to the second inference result M 209). According to the evaluation results corresponding to the multiple second inference results, a reference inference result 212 is determined from the multiple second inference results.
[0064] Finally, based on the reference inference result 212, the second language model 207 can be trained to obtain the target language model 213.
[0065] Next, through method embodiments, the model training method provided in this application will be introduced in detail.
[0066] See Figure 3 , Figure 3 which is a schematic flowchart of the model training method provided in the embodiments of this application. This model training method is executed by a computer device. For ease of description, hereinafter, the execution subject of this model training method is Figure 1 the server 140 in the computer system shown in Figure 3 as an example for introduction. As
[0067] shown, this model training method includes the following steps:
[0068] S301: According to the first input data, through the first language model, determine multiple first inference results. The first inference result includes the first inference path of the first language model for the first input data and the first inference conclusion generated by the first language model based on the first input data.
[0068] The first input data refers to the input data of the first language model in the self-supervised learning stage. The first input data can specifically be any natural language data. For example, the first input data can be a question in a mathematical reasoning task, such as "Calculate 15 + 27", etc. Another example is that the first input data can be a question in a dialogue task, such as "What's the weather like today", or natural language data input in other tasks, etc. In this regard, this application does not specifically limit the first input data.
[0069] The first language model refers to the large language model to be trained. For example, the first language model can include, but is not limited to, the Generative Pre-trained Transformer (GPT) model, the Bidirectional Encoder Representations from Transformers (Bert) model, the Text-to-Text Transfer Transformer (T5) model, the Large Language Model Meta AI (LLaMA), etc. In this regard, this application does not specifically limit the model structure of the first language model.
[0070] Multiple first inference results refer to multiple output results obtained by processing the first input data through the first language model. Each first inference result includes the first inference path of the first language model for the first input data and the first inference conclusion generated by the first language model based on the first input data. That is, the first inference result includes the processing result obtained when the first language model processes the first input data, and also includes the parsing method for describing how the processing result is obtained.
[0071] The first inference path is used to indicate the inference process of the first language model for the first input data. That is, the various inference steps when the first language model processes the first input data can be characterized by the first inference path. The first inference conclusion is the processing result generated by the first language model based on the first input data. That is, the processing result finally generated when the first language model processes the first input data can be used as the first inference conclusion. Taking the first input data as "Calculate 15 + 27" as an example, the first inference path includes "Step 1: Split 15 and 27 into 10 + 5 and 20 + 7 respectively. Step 2: First calculate 10 + 20 = 30. Step 3: Then calculate 5 + 7 = 12. Step 4: Finally, add 30 and 12 to get 42". The first inference conclusion is 42.
[0072] It should be understood that in practical applications, the same first input data can be input into the first language model multiple times, and the first language model will output different first inference results each time. Thus, for the same first input data, multiple first inference results corresponding to it can be obtained.
[0073] Exemplarily, the first input data can be obtained from a database, or alternatively, the first input data can be crawled from the network. After obtaining the first input data, the first input data can be input into the first language model to obtain multiple first inference results. That is, the first input data can be used as the input data of the first language model, and the first language model can process this input data multiple times to obtain multiple possible solution methods, that is, multiple first inference results can be obtained. For example, taking the first input data as "What is the quality of the current news", the first language model can analyze the quality of the current news from different dimensions to obtain multiple first inference results including the following processing results: the current news is a clickbait, the current news has vulgar pictures, the current news has no substantial content, etc., and each first inference result also includes a corresponding first inference path to indicate the analysis process for determining the above processing results.
[0074] S302: Determine a reasonable inference result among multiple first inference results according to a preset discrimination rule.
[0075] The preset discrimination rule refers to a rule for screening reasonable inference results, which is used to judge whether the first inference result is accurate and reasonable from a preset dimension. For example, the preset discrimination rule can be to judge whether the first inference path and the first inference conclusion in the first inference result match. If they match, the first inference result can be determined as a reasonable inference result. In this regard, the present application does not specifically limit the preset discrimination rule.
[0076] The reasonable inference result refers to the first inference result with relatively high accuracy among multiple first inference results, that is, the reasonable inference result is the first inference result that is closer to the correct inference result among multiple first inference results. In practical applications, one or more reasonable inference results can be determined from multiple first inference results. In this regard, the present application does not specifically limit the number of reasonable inference results.
[0077] S303: Train the first language model based on the reasonable inference result to obtain a second language model.
[0078] The second language model refers to a large language model obtained by training the first language model in the self-supervised learning stage. Its model structure is the same as that of the first language model, except that the model parameters are different.
[0079] After determining the reasonable inference result, the first language model can be trained based on the reasonable inference result. For example, the loss value during the model training process can be calculated based on the generation probability of the first inference path in the reasonable inference result. After determining the loss value, the first language model can be trained based on this loss value, that is, with the goal of maximizing the generation probability of the first inference path in the reasonable inference result, the model parameters of the first language model are adjusted to optimize the inference performance of the first language model to obtain the second language model. That is to say, the first language model after the training is completed can be used as the second language model.
[0080] It should be understood that in practical applications, the training end condition can be preset in advance. When the training of the first language model reaches this training end condition, the training of the first language model can be ended. The training end condition can be, for example, that the number of training rounds of the first language model reaches the preset round threshold. Or, for example, the performance of the first language model is tested and it is found that the performance of the first language model reaches the preset performance standard (such as reaching the preset inference accuracy, etc.). Or, for example, the performance of the first language model is tested and it is found that the performance of the first language model no longer improves significantly as the training progresses. The embodiments of the present application do not make any limitation on this training end condition.
[0081] S304: According to the second input data, through the second language model, determine multiple second inference results. The second inference result includes the second inference path of the second language model for the second input data and the second inference conclusion generated by the second language model based on the second input data. The second input data refers to the input data of the second language model in the reinforcement learning stage. The second input data can specifically be any natural language data. The second input data can be the same as or different from the first input data. In this regard, the present application does not specifically limit the second input data.
[0082] The second inference path is used to indicate the inference process of the second language model for the second input data. That is, the second inference path can be used to represent each inference step when the second language model processes the second input data. The second inference conclusion is the processing result generated by the second language model based on the second input data. That is, the final processing result generated when the second language model processes the second input data can be used as the second inference conclusion. Multiple second inference results refer to multiple output results obtained by processing the second input data through the second language model. Each second processing result includes the second inference path of the second language model for the second input data and the second inference conclusion generated by the second language model based on the second input data. That is, the second inference result includes the processing result obtained when the second language model processes the second input data, and also includes the parsing method for describing how the processing result is obtained.
[0083] It should be understood that in practical applications, the same second input data can be input into the second language model multiple times, and the second language model will output different second inference results each time. Therefore, for the same second input data, multiple second inference results corresponding to it can be obtained. In practical applications, in order to promote the generation of diverse second inference results, a temperature parameter T can be introduced during the training of the second language model. By setting the temperature parameter to be relatively large, the second inference results generated by the second language model based on the same second input data can be made more diverse, that is, there are larger differences among the multiple second inference results generated by the second language model. In addition, the strategy sampling method can also be used to make there be larger differences among the multiple second inference results generated by the second language model.
[0084] Exemplarily, the second input data can be obtained from a database, or alternatively, the second input data can be crawled from the network. After obtaining the second input data, the second input data can be input into the second language model to obtain multiple second inference results. That is, the second input data can be used as the input data of the second language model, and the second language model can process this input data multiple times to obtain multiple possible solution methods, that is, multiple second inference results can be obtained.
[0085] S305: Determine the evaluation results corresponding to each of the multiple second inference results.
[0086] The evaluation results are used to indicate the quality corresponding to the second inference results, and the evaluation results can include but are not limited to scores, evaluation grades, etc. In this regard, the present application does not specifically limit the manifestation form of the evaluation results.
[0087] Exemplarily, the evaluation results corresponding to each of the multiple second inference results can be determined through an evaluation model. For example, the multiple second inference results can be input into the evaluation model for processing, and then the evaluation results corresponding to each of the multiple second inference results can be obtained. Among them, the evaluation model refers to a model used to evaluate the quality corresponding to each of the multiple second inference results, and the evaluation model can include but are not limited to a state transition model, a reward model, and an inverse reinforcement learning model (IRL), etc. In this regard, the present application does not specifically limit the evaluation model.
[0088] Of course, in practical applications, the evaluation results corresponding to the second inference results can also be determined by other means. For example, a preset scoring mechanism can be used to evaluate the second inference results to obtain the corresponding evaluation results. The present application also does not specifically limit the method used to determine the evaluation results corresponding to the second inference results.
[0089] S306: Determine a reference inference result from multiple second inference results according to the evaluation results corresponding to the multiple second inference results.
[0090] The reference inference result refers to the second inference result with reference value among the multiple second inference results, that is, the reference inference result can provide valuable information for the training of the second language model in the reinforcement learning stage. The evaluation results corresponding to the multiple second inference results are used to indicate the quality corresponding to the multiple second inference results. Accordingly, a second inference result with high quality can be selected from the multiple second inference results as the reference inference result, or a second inference result with high quality and a second inference result with low quality can also be selected from the multiple second inference results as the reference inference result.
[0091] For example, taking the evaluation result as a score as an example, three second inference results with the highest scores can be selected from the scores corresponding to the multiple second inference results as the reference inference results. In this regard, the present application does not specifically limit the method for determining the reference inference result.
[0092] S307: Train the second language model based on the reference inference result to obtain the target language model.
[0093] After determining the reference inference result, the second language model can be trained with the reference inference result as feedback, so that the second language model can optimize the generation strategy to improve the inference ability and generation accuracy of the second language model.
[0094] The target language model refers to the large language model obtained after training the second language model in the reinforcement learning stage. The target language model can be applied to complex tasks. For example, the quality of news can be judged through the target language model. In this regard, the present application does not specifically limit the application scenario of the target language model.
[0095] Exemplarily, taking the reference inference result as a second inference result with high quality as an example, the loss value during the training process can be calculated based on the generation probability of the second inference path in the reference inference result. After determining the loss value, the second language model can be trained based on the loss value, that is, with the goal of maximizing the generation probability of the second inference path in the reference inference result, adjusting the parameters of the second language model, and optimizing the inference performance of the second language model to obtain the target language model.
[0096] In the case of using the evaluation model to determine the evaluation result corresponding to the second inference result, the second language model and the evaluation model can also be alternately trained iteratively for multiple rounds, so that the second language model after the training ends can be used as the target language model.
[0097] It should be understood that in practical applications, the training end condition can be preset in advance. When the training of the second language model reaches this training end condition, the training of the second language model can be ended. For example, the training end condition can be that the number of training rounds of the second language model reaches a preset round threshold. For another example, it can be to test the performance of the second language model and find that the performance of the second language model reaches a preset performance standard (such as reaching a preset inference accuracy, etc.). For still another example, it can be to test the performance of the second language model and find that the performance of the second language model no longer improves significantly as the training progresses. The embodiments of the present application do not make any limitation on this training end condition here.
[0098] The model training method provided by the embodiments of this application innovatively proposes a mechanism for training an LLM by combining self-supervised learning and reinforcement learning. In the self-supervised learning stage, a large amount of unlabeled first input data can be obtained at low cost; then, through the first language model, multiple first inference results are determined based on the first input data; and according to a preset discrimination rule, relatively accurate reasonable inference results are screened out from the multiple first inference results; furthermore, based on the screened reasonable inference results, the first language model is trained to enable the first language model to initially learn how to reason reasonably based on the input data and improve the inference ability of the first language model. After the above self-supervised learning stage, a second language model with relatively excellent inference ability will be obtained, and this second language model will be used as the training object in the reinforcement learning stage. Thus, by making the training object in the reinforcement learning stage initially have relatively excellent performance, it helps to improve the training efficiency in the reinforcement learning stage. In the reinforcement learning stage, a large amount of unlabeled second input data can still be obtained at low cost; then, through the second language model, multiple second inference results are determined based on the second input data. Since the second language model is trained by self-supervised learning, the second inference results generated by this second language model generally have high quality; then, evaluation results corresponding to the multiple second inference results are determined to give corresponding feedback information for the reinforcement learning stage through the evaluation results corresponding to the multiple second inference results; furthermore, based on the feedback information represented by the evaluation results corresponding to the multiple second inference results, reference inference results for participating in the reinforcement learning training are determined from the multiple second inference results, that is, from the relatively high-quality second inference results generated by the second language model trained in the self-supervised learning stage, reference inference results for guiding the model training direction in the reinforcement learning stage are further determined; finally, based on the reference inference results, the second language model is trained to enable the second language model to further learn how to reason more accurately and generate higher-quality inference results from the reference inference results, and deeply optimize the inference ability of this second language model to obtain the target language model. From the above self-supervised learning stage and reinforcement learning stage, it can be seen that this application can use a large amount of unlabeled data obtained at low cost to train the LLM, reduce the model training cost, and jointly train the LLM in two learning stages, which can rapidly and steadily improve the inference ability of the LLM.
[0099] · Determine reasonable inference results in the self-supervised learning stage and train the first language model accordingly
[0100] In a possible implementation manner, "determine reasonable inference results from multiple first inference results according to a preset discrimination rule" in S302 above may include S3021:
[0101] S3021: For each first inference result, determine the matching relationship between the first inference path and the first inference conclusion therein. If the matching relationship meets the preset matching requirements, determine the first inference result as a reasonable inference result.
[0102] When determining the matching relationship between the first inference path and the first inference conclusion in each first inference result, keywords can be extracted from the first inference path and the first inference conclusion respectively, and the semantic features corresponding to the first inference path and the first inference conclusion are determined. Furthermore, based on the matching relationship between the keywords in the first inference path and the keywords in the first inference conclusion, and the matching relationship between the semantic features of the first inference path and the semantic features of the first inference conclusion, the matching relationship between the first inference path and the first inference conclusion is determined.
[0103] For example, through Natural Language Processing (NLP) tools, word segmentation and part-of-speech tagging can be performed on the first inference path and the first inference conclusion respectively. Then, candidate keywords can be preliminarily screened from each segmented word according to the part-of-speech tagging results. Further, through the Term Frequency–-Inverse Document Frequency (TF-IDF) algorithm, the importance of each candidate keyword can be evaluated according to the word frequency and the inverse document frequency. Finally, the candidate keywords with importance higher than the preset importance threshold can be used as the keywords of the first inference path or the first inference conclusion.
[0104] At the same time, a pre-trained deep learning model, such as a pre-trained language representation model (Bidirectional Encoder Representations from Transformers, Bert), etc., can also be used to generate the corresponding semantic feature vectors based on the first inference path and the first inference conclusion. For example, each segmented word in the first inference path can be converted into a corresponding token, and then, according to the arrangement order of each segmented word in the first inference path, the converted tokens can be arranged into a corresponding token sequence, and start and end indicators are added at the front and end of the token sequence respectively; furthermore, the above token sequence is input into the Bert model for processing, so as to obtain the feature vector containing multiple layers output by the model, and the feature vector of the last layer output by the model can be used as the semantic feature of the first inference path. Similarly, the semantic feature of the first inference conclusion can be determined in the above manner.
[0105] After determining the keywords and semantic features corresponding to the first inference path and the first inference conclusion respectively, the matching relationship between the first inference path and the first inference conclusion can be determined based on the determined keywords and semantic features. For example, for each first inference result, it can be judged whether there are the same keywords or semantically matching keywords among the keywords corresponding to the first inference path and the keywords corresponding to the first inference conclusion. Additionally, the similarity algorithm can be used to calculate the similarity between the semantic features corresponding to the first inference path and the semantic features corresponding to the first inference conclusion. Thus, the matching relationship between the first inference path and the first inference conclusion is characterized by whether there are the same or semantically matching keywords between the first inference path and the first inference conclusion, and the semantic feature similarity between the first inference path and the first inference conclusion.
[0106] In the case where the matching relationship between the first inference path and the first inference conclusion in the first inference result is determined, it can be judged whether this matching relationship meets the preset matching requirements. The preset matching requirements refer to the requirements used to measure whether the first inference path and the first inference conclusion are reasonably matched. If the matching relationship meets the preset matching requirements, it is considered that the first inference path and the first inference conclusion in this first inference result are reasonably matched, and this first inference result can be used as a reasonable inference result. For example, if there are the same or semantically matching keywords between the first inference path and the first inference conclusion, and the semantic feature similarity between the first inference path and the first inference conclusion is greater than the preset semantic similarity threshold, it can be considered that the matching relationship between the first inference path and the first inference conclusion meets the preset matching requirements, and correspondingly, this first inference result can be used as a reasonable inference result; conversely, if there are no the same or semantically matching keywords between the first inference path and the first inference conclusion, or the semantic feature similarity between the first inference path and the first inference conclusion is less than or equal to the preset semantic similarity threshold, it can be considered that the matching relationship between the first inference path and the first inference conclusion does not meet the preset matching requirements, and correspondingly, this first inference result cannot be used as a reasonable inference result.
[0107] Thus, from the above content, it can be seen that the first inference results in which the first inference path and the first inference conclusion included are reasonably matched can be selected from multiple first inference results as the reasonable inference results for reference in the self-supervised learning stage, excluding the first inference results that are logically incoherent or whose conclusions do not match the inference steps, thereby ensuring that the determined reasonable inference results are relatively accurate and can reliably guide the first language model to learn how to perform inferences accurately during the self-supervised learning process, which is conducive to improving the inference ability of this first language model.
[0108] In a possible implementation, the first inference path includes at least one first inference step; the "training the first language model based on the reasonable inference result" in S303 may include S3031 to S3032:
[0109] S3031: Obtain the generation probability of each first inference step in the reasonable inference result, where the generation probability is determined according to the prediction probability of each token in the first inference step.
[0110] The first inference step refers to the step in the process of the first language model processing the first input data, and the first inference path is composed of at least one first inference step.
[0111] The generation probability of the first inference step is used to indicate the probability that the first language model generates this first inference step in the process of processing the first input data, and it is determined according to the prediction probability of each token in the first inference step. Each token in the first inference step refers to the smallest unit that makes up the first inference step, and the token can be a character, a segmented word, a symbol, etc.
[0112] When the first language model generates the first inference step in the process of processing the first input data, it generates the tokens in the first inference step one by one. When generating each token in the first inference step, the first language model first determines the respective original prediction scores (logits) corresponding to each candidate token, and converts the logits into the respective prediction probabilities corresponding to each candidate token through an activation function (softmax). The first language model can use the token with the highest corresponding prediction probability as the token in the first inference step.
[0113] After determining the tokens in a single first inference step in the above manner, that is, after determining a single first inference step, the prediction probabilities corresponding to the respective tokens in the first inference step can be obtained, and the generation probability of the first inference step can be determined based on these prediction probabilities. For example, the average probability obtained by averaging the prediction probabilities corresponding to the respective tokens in the first inference step can be used as the generation probability of the first inference step, or the sum probability obtained by summing the prediction probabilities corresponding to the respective tokens in the first inference step can be used as the generation probability of the first inference step. In this way, according to the above method, the generation probability corresponding to each first inference step in the reasonable inference result can be determined.
[0114] S3032: Train the first language model with the goal of maximizing the generation probabilities of the respective first inference steps in the reasonable inference result.
[0115] After determining the generation probability corresponding to each first inference step in the reasonable inference result, the first language model can be trained with the goal of maximizing the generation probability of each first inference step in the reasonable inference result.
[0116] Exemplarily, the loss value during the training process can be calculated through the generation probability of each first inference step in the reasonable inference result. For example, taking each first inference step included in the reasonable inference result as {Z1, Z2, …, Z n}, Z i represents the i-th first inference step. The loss value during the training process can be calculated through the loss function, and the formula of the loss function can refer to Formula 1.
[0117]
[0118] Among them, p(z i |x) represents the generation probability of the i-th first inference step under the condition of the given first input data x, L self-supervised represents the loss value in the self-supervised learning stage, and n is the number of first inference steps included in the reasonable inference result.
[0119] After determining the loss value, the first language model can be trained based on this loss value, that is, with the goal of maximizing the generation probability of each first inference step in the reasonable inference result, or equivalently, with the goal of minimizing this loss value. By continuously adjusting the parameters of the first language model, the inference performance of the first language model is optimized to obtain the second language model.
[0120] Thus, through the above content, it can be seen that after determining the generation probability of each first inference step in the reasonable inference result, the first language model is trained with the goal of maximizing the generation probability of each first inference step in the reasonable inference result. By maximizing the generation probability of each first inference step in the reasonable inference result, the first language model tends to learn how to generate reasonable inference results and learn how to make each inference step in the generated inference results as reasonable as possible, thereby improving the inference ability of the first language model. Moreover, in the self-supervised learning stage, by setting the above self-supervised task, the first language model can fully learn the reasonable inference ability without relying on manually labeled data and reducing the data acquisition cost.
[0121] · Determine the evaluation result of the second inference result in the reinforcement learning stage
[0122] In a possible implementation manner, "determine the evaluation result corresponding to each of the multiple second inference results" in S305 above may include S3051:
[0123] S3051: Based on the second inference paths and second inference conclusions respectively included in multiple second inference results, determine the evaluation scores respectively corresponding to the multiple second inference results through an evaluation model.
[0124] The evaluation scores respectively corresponding to the multiple second inference results refer to the evaluation results obtained by inputting the multiple second inference results into the evaluation model for quality evaluation.
[0125] After determining the second inference paths and second inference conclusions respectively included in the multiple second inference results, they can be input into the evaluation model for quality evaluation, and the evaluation scores respectively corresponding to the multiple second inference results can be obtained. That is, the quality respectively corresponding to the multiple second inference results can be characterized by the evaluation scores.
[0126] The evaluation model is used to conduct quality evaluation on the input inference results from a preset dimension. The preset dimension refers to the measurement criterion when the evaluation model conducts quality evaluation. The preset dimension may include at least one of an accuracy dimension, a diversity dimension, or an innovation dimension. The accuracy dimension means measuring the quality of the second inference result by accuracy, that is, whether the second inference result matches the preset correct inference result. The diversity dimension means measuring the quality of the second inference result by the diversity of the second inference result. For example, the diversity of the second inference result can be measured by comparing the similarity between a certain second inference conclusion and other second inference conclusions to avoid generating similar inference results. The innovation dimension means measuring the quality of the second inference result by the innovation of the second inference result. For example, the innovation of the second inference result can be measured by comparing the occurrence frequency of the second inference result to determine whether each second inference result can demonstrate innovation and new inference ideas.
[0127] Exemplarily, when conducting quality evaluation on the second inference result through the evaluation model, the second inference path and the second inference conclusion respectively included in each second inference result can be input separately, so that the evaluation model can conduct quality evaluation on the input second inference result from the accuracy dimension. That is, the evaluation model can measure whether the input second inference result is correct based on its own ability to evaluate whether the inference path and inference conclusion are correct, thereby obtaining the evaluation score of the second inference result. That is, the closer the second inference result is to the correct inference path and inference conclusion, the higher the evaluation score, and vice versa, the lower the evaluation score.
[0128] Alternatively, when conducting quality evaluation on the second inference result through the evaluation model, the second inference paths and second inference conclusions respectively included in multiple second inference results can also be input simultaneously, so that the evaluation model can comprehensively judge the evaluation scores respectively corresponding to the multiple second inference results from the accuracy dimension, the diversity dimension, and the innovation dimension. Refer to Figure 4 , Figure 4A schematic diagram for determining the evaluation scores corresponding to various second inference results provided by an embodiment of this application. Various second inference results (such as Figure 4 the second inference result 1 401,..., the second inference result M 402 in
[0129] can be input into the evaluation model 403. The evaluation model 403 can comprehensively evaluate the evaluation scores corresponding to various second inference results from the dimensions of accuracy, diversity, and innovation, and obtain the evaluation score 404 corresponding to the second inference result 1,..., the evaluation score 405 corresponding to the second inference result M.
[0130] For example, the ability of the evaluation model itself to evaluate whether the inference path and inference conclusion are correct can be used to determine the accuracy measurement result for the second inference result. The diversity measurement result for the second inference result can also be determined by comparing the similarities between the second inference results. The innovation measurement result for the second inference result can also be determined by comparing the repetition rates of the second inference paths in each second inference result. Finally, the evaluation scores corresponding to each second inference result can be determined by combining the above accuracy measurement result, diversity measurement result, and innovation measurement result. In this regard, this application does not specifically limit the number of second inference results input into the evaluation model and the specific working principle inside the evaluation model.
[0131] In a possible implementation manner, the evaluation model in S3051 above can initially be trained through S11 to S13:
[0132] S11: Obtain evaluation training samples. The evaluation training samples include various training inference results and their corresponding labeled scores. The training inference results include training inference paths and training inference conclusions.
[0133] The evaluation training samples refer to the training data for evaluating the model. The evaluation training samples include multiple training inference results and their respective corresponding annotation scores. The multiple training inference results refer to the input data of the evaluation model to be trained, and their respective corresponding annotation scores refer to the standard scores corresponding to the multiple training inference results, that is, the labels during the training process. The training inference results include training inference paths and training inference conclusions. That is, during the training process of the evaluation model, the training inference paths and training inference conclusions corresponding to the multiple training inference results can be input into the evaluation model to be trained for processing. The training inference path is used to indicate the parsing method of the second language model for the training input data. That is, the parsing process when processing the training input data can be characterized by the training inference path. Accordingly, the training inference conclusion is the processing result generated based on the training input data.
[0134] Exemplarily, the training inference results can be obtained from a database or crawled from the network. In this regard, the present application does not specifically limit the acquisition method of the training inference results. The annotation scores corresponding to the training inference results can be manually annotated.
[0135] S12: According to the training inference paths and training inference conclusions included in the multiple training inference results, through the evaluation model to be trained, determine the prediction scores corresponding to the multiple training inference results.
[0136] The prediction scores corresponding to the multiple training inference results refer to the prediction evaluation results obtained by inputting the multiple training inference results into the evaluation model to be trained for quality evaluation.
[0137] After obtaining the evaluation training samples, the training inference paths and training inference conclusions included in the multiple training inference results can be input into the evaluation model to be trained for processing. The evaluation model to be trained can perform quality evaluation on the input training inference results from a preset dimension, where the preset dimension includes at least one of the accuracy dimension, the diversity dimension, or the innovation dimension. Finally, the prediction scores corresponding to the multiple training inference results output by the evaluation model to be trained can be obtained.
[0138] Reference can be made to Figure 5 , Figure 5 which is a schematic diagram of the training evaluation model provided by the embodiments of the present application. After obtaining the evaluation training book 501, multiple training inference results (such as Figure 5The training inference results 1502, …, the training inference result K503) are input into the evaluation model 504 to be trained. The evaluation model 504 can comprehensively determine the prediction scores corresponding to various training inference results from the dimensions of accuracy, diversity, and innovation, and obtain the prediction score 505 corresponding to the training inference result 1, …, the prediction score 506 corresponding to the training inference result K. Then, the evaluation model 504 can be trained based on the prediction scores and annotation scores corresponding to various training inference results, that is, based on the difference between the prediction score 505 corresponding to the training inference result 1 and the annotation score 507 corresponding to the training inference result 1, and the difference between the prediction score 506 corresponding to the training inference result K and the annotation score 508 corresponding to the training inference result K, the evaluation model 504 is trained.
[0139] It should be noted that in the training process, the evaluation model to be trained evaluates the quality from which dimensions, and correspondingly, the evaluation model also evaluates the quality from the same dimensions during the application process.
[0140] S13: Train the evaluation model according to the prediction scores and annotation scores corresponding to various training inference results.
[0141] After determining the prediction scores and annotation scores corresponding to various training inference results, the difference between the prediction score and the annotation score can be calculated through a loss function, so that the evaluation model can be trained based on the calculated loss value.
[0142] Exemplarily, the loss function adopted during the training process of the evaluation model can be the mean square error function. Specifically, reference can be made to Formula 2.
[0143]
[0144] Among them, MSE represents the calculated loss value, n represents the number of training inference results, y i represents the annotation score corresponding to the i-th training inference result, r(x i ) represents the prediction score corresponding to the i-th training inference result, x i represents the i-th training inference result, and r() represents the evaluation model. In this regard, the present application does not specifically limit the loss function adopted during the training process of the evaluation model.
[0145] After calculating the loss value based on the loss function, the evaluation model can be trained based on this loss value, that is, aiming to reduce the loss value, by continuously adjusting the parameters of the evaluation model, optimizing the evaluation performance of the evaluation model, and improving the accuracy and stability of the evaluation model.
[0146] It should be understood that in practical applications, the training end condition can be preset in advance. When the training of the evaluation model reaches this training end condition, the training of the evaluation model can be ended. For example, the training end condition can be that the number of training rounds of the evaluation model reaches a preset round threshold. For another example, it can be to test the performance of the evaluation model and find that the performance of the evaluation model reaches a preset performance standard (such as reaching a preset evaluation accuracy, etc.). For yet another example, it can be to test the performance of the evaluation model and find that the performance of the evaluation model no longer improves significantly as the training progresses. The embodiments of the present application do not make any limitations on this training end condition here.
[0147] Thus, from the above content, it can be seen that the supervised learning method can be adopted to initially train the evaluation model through a large number of labeled evaluation training samples, so that the evaluation model can learn how to evaluate the inference results based on various training inference results and their respective corresponding labeled scores, thereby enabling the evaluation model to have a relatively high evaluation accuracy initially.
[0148] · Determine positive data and negative data as reference inference results in the reinforcement learning stage, and train the second language model accordingly
[0149] In a possible implementation manner, the "determine the reference inference result among multiple second inference results according to the evaluation results corresponding to the multiple second inference results" in S306 above may include S3061 to S3062:
[0150] S3061: Among the multiple second inference results, determine the high-quality second inference results whose corresponding evaluation results meet the high-quality result conditions, and determine the low-quality second inference results whose corresponding evaluation results meet the low-quality result conditions.
[0151] The high-quality second inference result refers to the second inference result that is closer to the correct inference result among the multiple second inference results, that is, the second inference result with a higher accuracy (positive data). The high-quality result condition refers to the condition for determining the high-quality second inference result. That is, the second inference result whose corresponding evaluation result among the multiple second inference results meets the high-quality result condition can be determined as the high-quality second inference result.
[0152] For example, the high-quality result condition is the second inference result whose corresponding evaluation result ranks among the top 3. Taking the evaluation result as the evaluation score as an example, reference can be made to Figure 6 , Figure 6 is a schematic diagram for determining the reference inference result provided by the embodiments of the present application. The multiple second inference results can be sorted in descending order according to the corresponding evaluation scores. Furthermore, the top 3 second inference results with the highest corresponding evaluation scores (such as Figure 6As shown, the second inference results 1, 5, and 4 are used as high-quality second inference results 601. Alternatively, taking the evaluation result as the evaluation level as an example, various second inference results corresponding to a high evaluation level can be used as high-quality second inference results. In this regard, the present application does not specifically limit the evaluation results corresponding to each of the multiple second inference results.
[0153] The inferior second inference result refers to the second inference result among the multiple second inference results that is far from the correct inference result, that is, the second inference result with lower accuracy (negative data). The inferior result condition refers to the condition for determining the inferior second inference result, that is, the second inference result among the multiple second inference results whose corresponding evaluation result meets the inferior result condition can be determined as the inferior second inference result.
[0154] For example, the inferior result condition is the second inference result whose corresponding evaluation result ranks among the last 3. Taking the evaluation result as the evaluation score as an example, as Figure 6 shown, the second inference results corresponding to the last 3 evaluation scores (as Figure 6 shown, the second inference results 3, 2, and 8) can be used as the inferior second inference result 602. Alternatively, taking the evaluation result as the evaluation level as an example, various second inference results corresponding to a low evaluation level can be used as the inferior second inference result.
[0155] S3062: Use both the high-quality second inference result and the inferior second inference result as the reference inference result.
[0156] As Figure 6 shown, after determining the high-quality second inference result 601 and the inferior second inference result 602, both the high-quality second inference result and the inferior second inference result can be used as the reference inference result 603, so that the second language model can not only learn how to generate high-quality inference results based on the high-quality second inference result during the training process, but also learn how to avoid generating incorrect inference results based on the inferior second inference result.
[0157] Thus, from the above content, it can be seen that the high-quality second inference result can be screened out from the multiple second inference results through the high-quality result condition, and the inferior second inference result can be screened out from the multiple second inference results through the inferior result condition. Then, both the high-quality second inference result and the inferior second inference result can be used as the reference inference result, so that during the model training process, the high-quality feedback indicated by the high-quality second inference result can be used to promote the second language model to optimize the generation of correct inference results, and the inferior feedback indicated by the inferior second inference result can be used to help the second language model correct the incorrect inference process, thereby helping to use richer training samples to improve the accuracy and inference ability of the second language model simultaneously from both positive and negative dimensions.
[0158] In a possible implementation, "training the second language model based on the reference inference result" in S307 above may include S3071 to S3072:
[0159] S3071: Obtain the generation probability of the second inference path in the high-quality second inference result as the high-quality generation probability, and obtain the generation probability of the second inference path in the low-quality second inference result as the low-quality generation probability.
[0160] The high-quality second inference result includes a second inference path, and the second inference path includes at least one second inference step. The second inference step refers to the step in the process of the second language model processing the second input data, and the second inference path is composed of at least one second inference step.
[0161] The generation probability of the second inference path refers to the probability of the second language model generating the second inference path in the process of processing the second input data. Since the second inference path is composed of at least one second inference step, the generation probability of the second inference path can be determined based on the generation probabilities of the respective second inference steps included in the second inference path.
[0162] The generation probability of the second inference step is used to indicate the probability of the second language model generating the second inference step in the process of processing the second input data, and it is determined according to the prediction probabilities of each token in the second inference step. The determination method of the generation probability of the second inference step is similar to that of the generation probability of the first inference step, the difference being that the first language model is replaced by the second language model and the first input data is replaced by the second input data. For details, refer to the content about determining the generation probability of the first inference step in the above text, which will not be elaborated here.
[0163] After determining the generation probabilities corresponding to the respective second inference steps in the second inference path, the generation probability of the second inference path can be calculated based on the generation probabilities corresponding to the respective second inference steps in the second inference path. For example, an average processing can be performed based on the generation probabilities corresponding to the respective second inference steps included in the second inference path, and the obtained mean probability can be used as the generation probability of the second inference path. Alternatively, a summation processing can be performed on the generation probabilities corresponding to the respective second inference steps in the second inference path, and the obtained sum probability can be used as the generation probability of the second inference path. Further, the generation probability of the second inference path in the high-quality second inference result can be used as the high-quality generation probability.
[0164] It should be understood that the inferior second inference result also includes a second inference path, and the second inference path includes at least one second inference step. According to the above method, the generation probability of the second inference path in the inferior second inference result can be determined. Further, the generation probability of the second inference path in the inferior second inference result can be used as the inferior generation probability.
[0165] S3072: Train the second language model with the goal of maximizing the high-quality generation probability and minimizing the inferior generation probability.
[0166] After determining the high-quality generation probability and the inferior generation probability, the second language model can be trained with the goal of maximizing the high-quality generation probability and minimizing the inferior generation probability.
[0167] Exemplarily, the loss value during the training process can be calculated through the high-quality generation probability. For example, the loss value during the training process can be calculated through a loss function, and the formula of the loss function can refer to Formula 3.
[0168]
[0169] where y i represents the i-th high-quality second inference result, p(r(y i )|x) represents the high-quality generation probability of the i-th high-quality second inference result, and L eval represents the loss value obtained based on the high-quality generation probability in the reinforcement learning stage.
[0170] Correspondingly, the inferior generation probability can also be substituted into Formula 3 to calculate the loss value obtained based on the inferior generation probability. When substituting the inferior generation probability into Formula 3, y i represents the i-th inferior second inference result, p(r(y i )|x) represents the inferior generation probability of the i-th inferior second inference result, and L eval represents the loss value obtained based on the inferior generation probability in the reinforcement learning stage.
[0171] After determining the above loss values, the second language model can be trained based on the above loss values, that is, with the goal of maximizing the high-quality generation probability, that is, with the goal of minimizing the loss value obtained based on the high-quality generation probability, and, with the goal of minimizing the inferior generation probability, that is, with the goal of maximizing the loss value obtained based on the inferior generation probability. By continuously adjusting the parameters of the second language model, the inference performance of the second language model is optimized to obtain the target language model.
[0172] Therefore, from the above, it can be seen that in the reinforcement learning stage, the training objective can be to maximize the probability of generating high-quality results and minimize the probability of generating low-quality results, so that the second language model can not only learn how to optimize the generation of correct reasoning results, but also learn how to avoid generating incorrect reasoning results, thereby effectively improving the reasoning ability and adaptability of the second language model and significantly improving the performance of the second language model in complex reasoning tasks.
[0173] In a possible implementation manner, when the evaluation result is the evaluation score determined by the evaluation model for the second reasoning result, "training the second language model with the goal of maximizing the probability of generating high-quality results and minimizing the probability of generating low-quality results" in S3072 above may include S21 to S23:
[0174] S21: Determine the positive loss according to the evaluation score of the high-quality second reasoning result and the generation probability of the second reasoning path in the high-quality second reasoning result.
[0175] The positive loss refers to the loss value determined based on the generation probability of the second reasoning path in the high-quality second reasoning result.
[0176] For reference, Figure 7 , Figure 7 is a schematic diagram for training the second language model provided in the embodiments of the present application. Exemplarily, the positive loss 703 during the training process can be calculated through a loss function based on the evaluation score 701 of the high-quality second reasoning result and the generation probability 702 of the second reasoning path in the high-quality second reasoning result. The formula of the loss function can be referred to formula 4.
[0177]
[0178] Among them, L positive represents the positive loss, represents the evaluation score of the i-th high-quality second reasoning result. A higher evaluation score will result in a smaller loss, and a lower evaluation score will result in a larger loss. represents the i-th high-quality second reasoning result, represents the generation probability of the second reasoning path in the i-th high-quality second reasoning result, that is, the high-quality generation probability of the i-th high-quality second reasoning result.
[0179] S22: Determine the negative loss according to the evaluation score of the low-quality second reasoning result and the generation probability of the second reasoning path in the low-quality second reasoning result.
[0180] The negative loss refers to the loss value determined based on the generation probability of the second reasoning path in the low-quality second reasoning result.
[0181] Exemplarily, such as Figure 7As shown, the negative loss 706 during the training process can be calculated through a loss function based on the evaluation score 704 of the inferior second inference result and the generation probability 705 of the second inference path in the inferior second inference result. The formula of the loss function can refer to Formula 5.
[0182]
[0183] Among them, L negative represents the negative loss, represents the evaluation score of the i-th inferior second inference result. A lower evaluation score will result in a larger loss. represents the i-th inferior second inference result, represents the generation probability of the second inference path in the i-th inferior second inference result, that is, the inferior generation probability of the i-th inferior second inference result.
[0184] S23: Based on the positive loss and the negative loss, train the second language model. The positive loss is used to achieve the goal of maximizing the high-quality generation probability, and the negative loss is used to achieve the goal of minimizing the inferior generation probability.
[0185] As Figure 7 shown, after determining the positive loss 703 and the negative loss 706, the second language model 707 can be trained based on the positive loss 703 and the negative loss 706. Among them, the positive loss is used to achieve the goal of maximizing the high-quality generation probability, that is, with the goal of maximizing the high-quality generation probability, or equivalently, with the goal of minimizing the positive loss, train the second language model. At the same time, the negative loss is used to achieve the goal of minimizing the inferior generation probability, that is, with the goal of minimizing the inferior generation probability, or equivalently, with the goal of maximizing the negative loss, by continuously adjusting the parameters of the second language model, optimize the inference performance of the second language model to obtain the target language model.
[0186] Therefore, from the above content, it can be seen that during the training of the second language model, the positive loss can be constructed based on the evaluation score of the high-quality second inference result and the generation probability of the second inference path therein, and the negative loss can be constructed based on the evaluation score of the inferior second inference result and the generation probability of the second inference path therein. Furthermore, the second language model can be backpropagated based on the positive loss and the negative loss. By continuously adjusting the parameters of the second language model, optimize the performance of the second language model to accelerate model convergence, improve the inference ability of the second language model, so that the second language model can quickly optimize the inference strategy when facing complex tasks and avoid over-reliance on incorrect inference paths.
[0187] · Alternately train the second language model and the evaluation model in the reinforcement learning stage
[0188] In a possible implementation, "training the second language model based on the reference inference result to obtain the target language model" in S307 above may include:
[0189] Performing iterative multi-round alternating training on the second language model and the evaluation model, and using the second language model obtained in the last round of alternating training as the target language model; wherein, in each round of alternating training, the second language model is trained based on the reference inference result, and the evaluation model is trained based on multiple third inference results generated by the second language model.
[0190] Multi-round alternating training means performing multi-round alternating training on the second language model and the evaluation model, that is, in each round of alternating training, after the second language model is trained, the evaluation model is trained.
[0191] Exemplarily, after training the second language model based on the positive loss and the negative loss through S3071 and S3072 above, the evaluation model can be trained to complete one round of alternating training. That is, in the j-th (j is an integer greater than or equal to 1) round of alternating training, the second language model can be trained first, and then the evaluation model can be trained. In the (j + 1)-th round of alternating training, the second language model obtained in the j-th round of alternating training can be continuously trained, and the evaluation model obtained in the j-th round of alternating training can be continuously trained. In this way, by performing iterative multi-round alternating training, the second language model obtained in the last round of alternating training can be used as the target language model.
[0192] It should be noted that in the process of each round of alternating training, the second language model can be trained based on the reference inference result, that is, the second language model can be trained based on the high-quality second inference result and the low-quality second inference result. Specifically, the training steps of S21 - S23 above can be referred to. After the training of the second language model is completed, the evaluation model can be trained based on multiple third inference results generated by the second language model. Among them, the multiple third inference results generated by the second language model refer to multiple inference results generated by the second language model based on new input data.
[0193] In a possible implementation, "training the evaluation model based on multiple third inference results generated by the second language model" above may include S31 to S34:
[0194] S31: According to the third input data, determine multiple third inference results through the second language model, and determine the generation probability corresponding to each of the multiple third inference results.
[0195] The third input data refers to the input data of the second language model during the alternating training process. The third input data may be the same as the first input data or the second input data, or may not be the same as the first input data or the second input data. In this regard, the present application does not specifically limit the third input data.
[0196] Inputting the third input data into the second language model can obtain multiple third inference results.
[0197] The generation probability corresponding to each of the multiple third inference results refers to the probability of generating the third inference result during the process of the second language model processing the third input data. Each third inference result includes a third inference path and a third inference conclusion. The generation probability of the third inference path can be determined by referring to the steps in S3071 above, which will not be elaborated here. Correspondingly, the generation probability of the third inference conclusion can also be determined by referring to the steps for determining the generation probability of the inference path. After determining the generation probability of the third inference path and the generation probability of the third inference conclusion included in each third inference result, an averaging process can be performed on the generation probability of the third inference path and the generation probability of the third inference conclusion. The obtained mean probability can be used as the generation probability of this third inference result. In this way, according to the above method, the generation probability corresponding to each of the multiple third inference results can be determined.
[0198] S32: Determine the annotation scores corresponding to each of the multiple third inference results according to the generation probabilities corresponding to each of the multiple third inference results.
[0199] The annotation scores corresponding to each of the multiple third inference results refer to the well-annotated evaluation scores corresponding to each of the multiple third inference results, which can be used as the annotation data during the training process.
[0200] Exemplarily, a mapping relationship between the generation probability of the inference result and the annotation score can be preset. That is, based on this mapping relationship, according to the generation probabilities corresponding to each of the multiple third inference results, the corresponding annotation scores can be found from this mapping relationship. It should be understood that in the mapping relationship, the higher the generation probability, the higher the corresponding annotation score. In this regard, the present application does not specifically limit the manner of determining the annotation scores corresponding to each of the multiple third inference results.
[0201] S33: Determine the prediction scores corresponding to each of the multiple third inference results through an evaluation model according to the multiple third inference results.
[0202] The prediction scores corresponding to each of the multiple third inference results refer to the output results obtained by inputting the multiple third inference results into the evaluation model for quality evaluation. Inputting the multiple third inference results into the evaluation model for quality evaluation can obtain the prediction scores corresponding to each of the multiple third inference results.
[0203] S34: Train an evaluation model based on the prediction scores and annotation scores corresponding to various third inference results.
[0204] After obtaining the prediction scores and annotation scores corresponding to various third inference results, the difference between the prediction scores and the annotation scores can be calculated through a loss function, so that the evaluation model can be trained based on the calculated loss value.
[0205] Exemplarily, the loss function in Formula 2 can be used to calculate the loss value. In this regard, the present application does not specifically limit the loss function used in training the evaluation model. After calculating the loss value based on the loss function, the evaluation model can be trained based on this loss value, that is, aiming to reduce the loss value, and by continuously adjusting the parameters of the evaluation model, the evaluation performance of the evaluation model is optimized to improve the accuracy and stability of the evaluation model.
[0206] In this way, it can be seen from the above content that during the multi-round alternating training process, the evaluation model can be optimized based on the inference results generated by the second language model, so that through multi-round alternating training, the inference ability of the second language model and the evaluation accuracy of the evaluation model can be continuously improved alternately.
[0207] Therefore, it can be seen from the above content that in the reinforcement learning stage, iterative multi-round alternating training can be performed on the second language model and the evaluation model to continuously optimize the generation performance of the second language model and the evaluation performance of the evaluation model, promote the second language model to continuously improve the quality of the inference results, and the evaluation model to continuously improve its evaluation accuracy. Moreover, the process of jointly optimizing the two ensures the simultaneous improvement of the diversity, innovation, and accuracy of the inference results, thereby enhancing the inference ability and decision-making ability of the second language model.
[0208] · Alternately execute the self-supervised learning stage and the reinforcement learning stage
[0209] In a possible implementation manner, the method provided in the embodiments of the present application may further include S41:
[0210] S41: Among various second inference results, determine reasonable second inference results whose corresponding evaluation results meet the reasonable result conditions, and update the reasonable inference results using the reasonable second inference results.
[0211] The reasonable second inference result refers to the second inference result with relatively high accuracy among various second inference results, which can be used as the training data for a new round of self-supervised learning stage.
[0212] The reasonable result condition refers to the condition used to determine the reasonable second inference result, that is, among multiple second inference results, the second inference result corresponding to the evaluation result that meets the reasonable result condition can be determined as the reasonable second inference result. The reasonable result condition can, for example, be the same as the high-quality result condition introduced above, or it can be different from the high-quality result condition. In this regard, the present application does not specifically limit the reasonable result condition.
[0213] Exemplarily, the reasonable result condition is the second inference result whose corresponding evaluation result ranks among the top 5. Taking the evaluation result as the evaluation score as an example, multiple second inference results can be sorted in descending order according to the corresponding evaluation scores. Furthermore, the top 5 second inference results with the highest corresponding evaluation scores can be used as the reasonable second inference results.
[0214] After determining the reasonable second inference result, the reasonable inference result used as training data in the self-supervised learning stage can be updated using the reasonable second inference result. For example, the reasonable second inference result can be mixed with the reasonable inference result determined in the self-supervised learning stage, and a preset number of inference results can be randomly selected therefrom as the updated reasonable inference result.
[0215] Correspondingly, "training the second language model based on the reference inference result to obtain the target language model" in S307 above can include the following S42 and S43:
[0216] S42: Training the second language model based on the reference inference result to obtain the third language model.
[0217] The third language model refers to the language model obtained after being trained in the reinforcement learning stage, and it can be the language model obtained after multiple rounds of alternating training of the second language model and the evaluation model.
[0218] Specifically, the operation steps of training the second language model based on the reference inference result are the same as those of S3071 to S3072 above and will not be elaborated here.
[0219] S43: Training the third language model based on the updated reasonable inference result to obtain the target language model.
[0220] After training the third language model in the reinforcement learning phase, based on the updated reasonable inference results, self-supervised learning can be performed on the third language model to obtain the target language model. That is, self-supervised learning phase and reinforcement learning phase can be alternately executed. After completing the training in the reinforcement learning phase, the reasonable second inference results with higher quality in the reinforcement learning phase can be mixed with the reasonable inference results in the self-supervised learning phase, and the training in the new round of self-supervised learning phase can be continued, that is, the third language model can be trained based on the updated reasonable inference results. Among them, the operation steps of training the third language model based on the updated reasonable inference results are similar to the steps of S3031 to S3032 above, except that the reasonable inference results are replaced by the updated reasonable inference results, and the training object is replaced from the first language model to the third language model, which will not be elaborated here.
[0221] As an example, after completing the self-supervised learning of the third language model based on the updated reasonable inference results, the trained language model can be directly used as the target language model.
[0222] Thus, from the above content, it can be seen that in this application, not only can alternating training be performed within the reinforcement learning phase, but also alternating training can be performed based on the self-supervised learning phase and the reinforcement learning phase. The better second inference results generated in the reinforcement learning phase can be used to update the training data in the self-supervised learning phase. After completing the training in the reinforcement learning phase, self-supervised learning can be performed on the LLM trained in the reinforcement learning phase based on the updated reasonable inference results, so as to perform self-supervised learning again using the higher-quality training data generated in the reinforcement learning phase, which is beneficial to further improving the performance of the LLM.
[0223] In a possible implementation manner, the "training the third language model based on the updated reasonable inference results to obtain the target language model" in S43 above may include S51 to S55:
[0224] S51: Train the third language model based on the updated reasonable inference results to obtain the fourth language model.
[0225] The fourth language model refers to the large language model obtained by training the third language model in the new round of self-supervised learning phase. Its model structure is the same as that of the third language model, except that the model parameters are different.
[0226] S52: According to the fourth input data, determine a variety of fourth inference results through the fourth language model.
[0227] The fourth input data refers to the input data of the fourth language model in a new round of reinforcement learning stage. The fourth input data can specifically be any natural language data, and this fourth input data can be the same as or different from any one of the first input data, the second input data, and the third input data. In this regard, the present application does not specifically limit the fourth input data.
[0228] Multiple fourth inference results refer to multiple output results obtained by processing the fourth input data through the fourth language model. Each fourth inference result includes the inference path of the fourth language model for the fourth input data and the inference conclusion generated by the fourth language model based on the fourth input data.
[0229] It should be understood that in practical applications, the same fourth input data can be input into the fourth language model multiple times, and the fourth language model will output different fourth inference results each time. Thus, for the same fourth input data, multiple fourth inference results corresponding to it can be obtained.
[0230] S53: Determine the evaluation results corresponding to each of the multiple fourth inference results.
[0231] Exemplarily, the evaluation results corresponding to each of the multiple fourth inference results can be determined through an evaluation model. For example, the multiple fourth inference results can be input into the evaluation model for processing, and then the evaluation results corresponding to each of the multiple fourth inference results can be obtained. Specifically, the operation steps for determining the evaluation results corresponding to each of the multiple fourth inference results are similar to the steps of S3051 above, except that the multiple second inference results are replaced by the multiple fourth inference results, which will not be elaborated here.
[0232] S54: Select at least one fourth inference result from the multiple fourth inference results according to the evaluation results corresponding to each of the multiple fourth inference results, and update the reference inference result.
[0233] After determining the evaluation results corresponding to each of the multiple fourth inference results, at least one fourth inference result can be selected from the multiple fourth inference results as the new reference inference result. Then, it can be mixed with the reference inference result in the reinforcement learning stage of the previous round of alternating training to obtain the updated reference inference result.
[0234] Exemplarily, the methods introduced in S3061 to S3062 above can be adopted to select high-quality fourth inference results and low-quality fourth inference results from the multiple fourth inference results. Furthermore, the selected high-quality fourth inference results and low-quality fourth inference results are used to update the reference inference result.
[0235] S55: Train the fourth language model based on the updated reference inference result to obtain the target language model.
[0236] After obtaining the updated reference inference result, based on the updated reference inference result, the fourth language model can be trained, with the training result obtained in this round of reinforcement learning stage as the target language model. That is, after determining the updated reference inference result, the updated reference inference result can be used as feedback to train the fourth language model, so that the fourth language model can optimize the generation strategy to improve the inference ability and generation accuracy of the fourth language model. Furthermore, the fourth language model with better performance can be used as the finally required target language model.
[0237] Specifically, the operation steps of training the fourth language model based on the updated reference inference result are the same as those of S3071 - S3072 above, except that the training object is replaced by the fourth language model, which will not be elaborated here. Moreover, when training the fourth language model, the fourth language model and the evaluation model can also be alternately trained.
[0238] Thus, from the above content, it can be seen that in the alternating training of the self - supervised learning stage and the reinforcement learning stage, based on the updated reasonable inference result determined in the reinforcement learning stage, after training in the self - supervised learning stage, the training in the reinforcement learning stage can continue based on the fourth language model trained in the self - supervised learning stage, so that the language model trained in the reinforcement learning stage is used as the target language model. Therefore, by alternately performing the self - supervised learning stage and the reinforcement learning stage, the performance and accuracy of the LLM model can be continuously iteratively improved.
[0239] It should be understood that in practical applications, several rounds of self - supervised learning and reinforcement learning can be alternately executed according to actual needs, that is, the embodiments of the present application do not make any limitations on the number of rounds of self - supervised learning and reinforcement learning performed. In addition, according to actual needs, the language model trained through the last round of self - supervised learning can be selected as the target language model, or the language model trained through the last round of reinforcement learning can be selected as the target language model.
[0240] · Provide an overall exemplary introduction to the model training process
[0241] Finally, reference can be made to Figure 8 for an overall exemplary introduction to the embodiments of the present application. Figure 8 is the overall architecture schematic diagram of the model training method provided by the embodiments of the present application.
[0242] As shown in Figure 8As shown by 801 in [reference], during the self-supervised learning phase, the first input data can be input into the first language model to obtain multiple first inference results. Then, according to a preset discrimination rule, a reasonable inference result can be selected from them. After determining the reasonable inference result, the generation probability of each first inference step in the reasonable inference result can be obtained. Subsequently, with the goal of maximizing the generation probability of each first inference step in the reasonable inference result, the first language model is trained.
[0243] As Figure 8 As shown by 802 in [reference], during the reinforcement learning phase, the second input data is input into the second language model, and multiple second inference results can be obtained. The multiple second inference results are input into an evaluation model for quality evaluation, and evaluation results corresponding to each of the multiple second inference results can be obtained. Based on the good result conditions and bad result conditions, good second inference results and bad second inference results are determined from the multiple second inference results as reference inference results. Then, according to the evaluation score of the good second inference result and the generation probability of the second inference path in the good second inference result, a positive loss is determined. According to the evaluation score of the bad second inference result and the generation probability of the second inference path in the bad second inference result, a negative loss is determined. Finally, based on the positive loss and the negative loss, the second language model is trained.
[0244] Finally, as Figure 8 As shown by 803 in [reference], during the reinforcement learning phase, iterative multi-round alternating training can be performed on the second language model and the evaluation model, so that the inference strategy of the second language model can be adjusted based on the evaluation score of the evaluation model, and the scoring accuracy of the evaluation model can be optimized through the third inference result output by the second language model, thereby jointly improving the quality of the inference result output by the target language model.
[0245] To verify the performance of the target language model, after obtaining the target language model, the embodiments of the present application test it on multiple datasets. According to the tests, the accuracy rate of the target language model on the Mathematics Reasoning (MATH 500) test set has increased by 10%, the accuracy rate on the Code Reasoning (Codeforces) test set has increased by 12%, and the reasoning ability on the American Invitational Mathematics Examination (AIME) test set has increased by 8%. Thus, it can be seen that the target language model can optimize each reasoning step during the reasoning process, thereby improving the reasoning accuracy. Moreover, during the training process of the target language model, through joint training with high-quality second reasoning results and low-quality second reasoning results, the target language model can effectively learn experience from incorrect reasoning results, improving the robustness of the reasoning of the target language model. And through testing, it is found that the error rate of the target language model has been reduced by an average of 10%, and it can well adapt to complex reasoning tasks.
[0246] · Applying the target language model
[0247] After obtaining the target language model, the target language model can be applied to actual complex tasks. For example, the target language model can be applied to tasks such as news quality recognition.
[0248] In a possible implementation manner, the method provided by the embodiments of the present application may further include S308 to S309:
[0249] S308: Obtain the target news content.
[0250] The target news content refers to the input data of the target language model, that is, the news content to be classified. Through the target language model, the type of the target news content can be determined, that is, the category to which the target news content belongs can be determined through the target language model. For example, the server can obtain the target news content from the database.
[0251] S309: According to the target news content, through the target language model, determine the category to which the target news content belongs.
[0252] After obtaining the target news content, the target news content can be input into the target language model for processing by the target language model, so as to obtain the category to which the target news content belongs and the processing process of how to obtain the category to which the target news content belongs. Among them, the category to which the target news content belongs is used to indicate the quality category of the target news content. For example, when the target news content is determined to be low-quality news by the target language model, the category to which the target news content belongs can specifically be vulgar category, clickbait category, and lack of substantial content category, etc. In this regard, the present application does not specifically limit the category to which the target news content belongs.
[0253] In addition, reference can be made to Figure 9 , Figure 9 which is a schematic diagram of the news low-quality residue rate provided by the embodiments of the present application. By processing the target news content through the target language model, different categories of low-quality news can be more accurately identified, thereby effectively reducing the generation of low-quality residual data. As shown by the full-scene low-quality residue rate in Figure 9 , after applying the target language model, the low-quality residue rate has decreased significantly (from 3.53% to 0.30%).
[0254] Therefore, from the above content, it can be seen that after obtaining the target language model, it is possible to identify whether the target news content is low-quality news through the target language model, and moreover, the category of low-quality news can be accurately identified, thereby improving the category recognition efficiency and the recommendation quality and user experience of news content.
[0255] In addition, the target language model trained through the embodiments of the present application can also be applied to other natural language tasks. For example, it can be applied to the task of extracting text summaries. In this task, the user can input a long document into the target language model to generate a summary through the target language model, and a summary result including the summary extraction process can be obtained. Another example is that it can be applied to the translation task. In this task, the user can input the material to be translated into the target language model to perform translation through the target language model, and a translation result including the translation process can be obtained. Another example is that it can be applied to the dialogue task. In this task, the user can input the dialogue text into the target language model to obtain the corresponding reply content through the target language model, and the parsing process of the reply text based on the input dialogue text. In this regard, the present application does not specifically limit the application scenarios of the target language model.
[0256] Based on the model training method provided in the foregoing embodiments, the present application also correspondingly provides a model training device. The following will be described in conjunction with Figure 10 as follows. Figure 10 which is a schematic structural diagram of the model training device 1000 provided by the embodiments of the present application. The device includes:
[0257] The first inference module 1001 is configured to determine multiple first inference results through a first language model according to first input data, where the first inference results include a first inference path of the first language model for the first input data and a first inference conclusion generated by the first language model based on the first input data;
[0258] The first screening module 1002 is configured to determine reasonable inference results from the multiple first inference results according to a preset discrimination rule;
[0259] The first training module 1003 is configured to train the first language model based on the reasonable inference results to obtain a second language model;
[0260] The second inference module 1004 is configured to determine multiple second inference results through the second language model according to second input data, where the second inference results include a second inference path of the second language model for the second input data and a second inference conclusion generated by the second language model based on the second input data;
[0261] The evaluation module 1005 is configured to determine evaluation results corresponding to the multiple second inference results respectively;
[0262] The second screening module 1006 is configured to determine reference inference results from the multiple second inference results according to the evaluation results corresponding to the multiple second inference results respectively;
[0263] The second training module 1007 is configured to train the second language model based on the reference inference results to obtain a target language model.
[0264] Optionally, the first screening module 1002 is specifically configured to:
[0265] For each of the first inference results, determine the matching relationship between the first inference path and the first inference conclusion therein. If the matching relationship meets the preset matching requirements, determine the first inference result as the reasonable inference result.
[0266] Optionally, the first inference path includes at least one first inference step; the first training module 1003 is specifically configured to:
[0267] Obtain the generation probability of each first inference step in the reasonable inference results, where the generation probability is determined according to the prediction probability of each token in the first inference step;
[0268] Take maximizing the generation probability of each first inference step in the reasonable inference results as the objective to train the first language model.
[0269] Optionally, the evaluation module 1005 is specifically configured to:
[0270] According to the second inference paths and the second inference conclusions included in the multiple second inference results, determine the evaluation scores corresponding to the multiple second inference results through an evaluation model;
[0271] Wherein, the evaluation model is used to evaluate the quality of the input inference results from a preset dimension, and the preset dimension includes at least one of an accuracy dimension, a diversity dimension, or an innovation dimension.
[0272] Optionally, the device further includes the following modules for training the initial evaluation model:
[0273] A first acquisition module, configured to acquire evaluation training samples, where the evaluation training samples include multiple training inference results and their respective corresponding labeled scores, and the training inference results include training inference paths and training inference conclusions;
[0274] A first determination module, configured to determine the prediction scores corresponding to the multiple training inference results through the evaluation model to be trained according to the training inference paths and the training inference conclusions included in the multiple training inference results;
[0275] A third training module, configured to train the evaluation model according to the prediction scores and the labeled scores corresponding to the multiple training inference results.
[0276] Optionally, the second screening module 1006 is specifically configured to:
[0277] Among the multiple second inference results, determine high-quality second inference results whose corresponding evaluation results meet the high-quality result conditions, and determine low-quality second inference results whose corresponding evaluation results meet the low-quality result conditions;
[0278] Use both the high-quality second inference results and the low-quality second inference results as the reference inference results.
[0279] Optionally, the second training module 1007 is specifically configured to:
[0280] Obtain the generation probability of the second inference path in the high-quality second inference result as the high-quality generation probability, and obtain the generation probability of the second inference path in the low-quality second inference result as the low-quality generation probability;
[0281] Train the second language model with the goal of maximizing the high-quality generation probability and minimizing the low-quality generation probability.
[0282] Optionally, the evaluation result is an evaluation score determined by an evaluation model for the second inference result; the second training module 1007 is specifically configured to:
[0283] Determine a positive loss according to the evaluation score of the high-quality second inference result and the generation probability of the second inference path in the high-quality second inference result;
[0284] Determine a negative loss according to the evaluation score of the low-quality second inference result and the generation probability of the second inference path in the low-quality second inference result;
[0285] Train the second language model based on the positive loss and the negative loss, where the positive loss is used to achieve the goal of maximizing the high-quality generation probability, and the negative loss is used to achieve the goal of minimizing the low-quality generation probability.
[0286] Optionally, the second training module 1007 is specifically configured to:
[0287] Perform iterative multi-round alternating training on the second language model and the evaluation model, and use the second language model obtained in the last round of the alternating training as the target language model;
[0288] Wherein, in each round of the alternating training, the second language model is trained based on the reference inference result, and the evaluation model is trained based on multiple third inference results generated by the second language model.
[0289] Optionally, the second training module 1007 is specifically configured to:
[0290] Determine the multiple third inference results through the second language model according to the third input data, and determine the generation probability corresponding to each of the multiple third inference results;
[0291] Determine the annotation score corresponding to each of the multiple third inference results according to the generation probability corresponding to each of the multiple third inference results;
[0292] Determine the prediction score corresponding to each of the multiple third inference results through the evaluation model according to the multiple third inference results;
[0293] Train the evaluation model according to the prediction score and the annotation score corresponding to each of the multiple third inference results.
[0294] Optionally, the device further includes:
[0295] A first update module, configured to determine, from the multiple second inference results, reasonable second inference results whose corresponding evaluation results meet the reasonable result condition, and update the reasonable inference results by using the reasonable second inference results;
[0296] Correspondingly, the second training module 1007 is specifically configured to:
[0297] Train the second language model based on the reference inference result to obtain a third language model;
[0298] Train the third language model based on the updated reasonable inference result to obtain the target language model.
[0299] Optionally, the second training module 1007 is specifically configured to:
[0300] Train the third language model based on the updated reasonable inference result to obtain a fourth language model;
[0301] Determine multiple fourth inference results through the fourth language model according to the fourth input data;
[0302] Determine the evaluation results corresponding to the multiple fourth inference results respectively;
[0303] Select at least one of the fourth inference results from the multiple fourth inference results according to the evaluation results corresponding to the multiple fourth inference results respectively, and update the reference inference result;
[0304] Train the fourth language model based on the updated reference inference result to obtain the target language model.
[0305] An embodiment of the present application further provides a computer device, which may specifically be a terminal device or a server. The terminal device and the server provided by the embodiment of the present application will be introduced from the perspective of hardware implementation below.
[0306] See Figure 11 , Figure 11 is a schematic structural diagram of the terminal device provided by the embodiment of the present application. As Figure 11 shown, for the sake of convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The terminal may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), an in-vehicle computer, etc. Taking the terminal as a computer as an example:
[0307] Figure 11The block diagram of a part of the structure of a computer related to the terminal provided in the embodiment of the present application is shown. Refer to Figure 11 , the computer includes: a Radio Frequency (RF) circuit 1210, a memory 1220, an input unit 1230 (including a touch panel 1231 and other input devices 1232), a display unit 1240 (including a display panel 1241), a sensor 1250, an audio circuit 1260 (connected with a speaker 1261 and a microphone 1262), a wireless fidelity (WiFi) module 1270, a processor 1280, and a power supply 1290, etc. Those skilled in the art can understand that Figure 11 the computer structure shown in
[0308] does not limit the computer, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0309] The processor 1280 is the control center of the computer, connecting various parts of the entire computer through various interfaces and lines, and executing various functions of the computer and processing data by running or executing the software programs and / or modules stored in the memory 1220, and calling the data stored in the memory 1220. Optionally, the processor 1280 may include one or more processing units; preferably, the processor 1280 may integrate an application processor and a modem processor, where the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1280.
[0310] In the embodiment of the present application, the processor 1280 included in the terminal is used to execute the steps in the model training method described in each of the foregoing embodiments.
[0311] See Figure 12 , Figure 12Schematic diagram of a server 1300 provided by an embodiment of the present application. The server 1300 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1322 (for example, one or more processors) and a memory 1332, and one or more storage media 1330 (for example, one or more mass storage devices) for storing application programs 1342 or data 1344. Among them, the memory 1332 and the storage media 1330 may be transient storage or persistent storage. The program stored in the storage media 1330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 1322 may be configured to communicate with the storage media 1330 and execute a series of instruction operations in the storage media 1330 on the server 1300.
[0312] The server 1300 may further include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input / output interfaces 1358, and / or, one or more operating systems, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0313] The steps performed by the server in the above embodiments may be based on the Figure 12 server structure shown. Among them, the CPU 1322 is used to execute the steps in the model training method described in each of the foregoing embodiments.
[0314] An embodiment of the present application further provides a computer-readable storage medium for storing a computer program, and the computer program is used to execute the steps in the model training method described in each of the foregoing embodiments.
[0315] An embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the model training method described in each of the foregoing embodiments.
[0316] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0317] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0318] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0319] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0320] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store computer programs.
[0321] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the relationship between associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item) of the following" or its similar expression refers to any combination of these items, including any combination of a single item or multiple items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0322] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.
[0323] As mentioned above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A model training method, characterized in that, The method includes: Based on first input data, determining multiple first inference results through a first language model, where the first inference results include a first inference path of the first language model for the first input data and a first inference conclusion generated by the first language model based on the first input data; Determining a reasonable inference result among the multiple first inference results according to a preset discrimination rule; Training the first language model based on the reasonable inference result to obtain a second language model; Based on second input data, determining multiple second inference results through the second language model, where the second inference results include a second inference path of the second language model for the second input data and a second inference conclusion generated by the second language model based on the second input data; Determining evaluation results corresponding to the multiple second inference results respectively; Determining a reference inference result among the multiple second inference results according to the evaluation results corresponding to the multiple second inference results respectively; Training the second language model based on the reference inference result to obtain a target language model.
2. The method according to claim 1, wherein The determining a reasonable inference result among the multiple first inference results according to a preset discrimination rule includes: For each of the first inference results, determining a matching relationship between the first inference path and the first inference conclusion therein. If the matching relationship meets a preset matching requirement, determining the first inference result as the reasonable inference result.
3. The method according to claim 1 or 2, characterized in that, The first inference path includes at least one first inference step; the training the first language model based on the reasonable inference result includes: Obtaining the generation probability of each first inference step in the reasonable inference result, where the generation probability is determined according to the prediction probability of each token in the first inference step; Training the first language model with the goal of maximizing the generation probability of each first inference step in the reasonable inference result.
4. The method according to any one of claims 1 to 3, characterized in that The determining the evaluation results corresponding to the multiple second inference results respectively includes: Determining evaluation scores corresponding to the multiple second inference results respectively through an evaluation model according to the second inference paths and the second inference conclusions included in the multiple second inference results respectively; Wherein, the evaluation model is used to evaluate the quality of the input inference result from a preset dimension, and the preset dimension includes at least one of an accuracy dimension, a diversity dimension or an innovation dimension.
5. The method according to claim 4, characterized in that, The initial evaluation model is trained in the following manner: Obtaining evaluation training samples, where the evaluation training samples include multiple training inference results and their respective labeled scores, and the training inference results include training inference paths and training inference conclusions; Determining prediction scores corresponding to the multiple training inference results respectively through the evaluation model to be trained according to the training inference paths and the training inference conclusions included in the multiple training inference results respectively; Training the evaluation model according to the prediction scores and the labeled scores corresponding to the multiple training inference results respectively.
6. The method according to any one of claims 1 to 5, characterized in that Determining a reference inference result from the multiple second inference results according to the evaluation results respectively corresponding to the multiple second inference results includes: Among the multiple second inference results, determining high-quality second inference results whose corresponding evaluation results meet the high-quality result conditions, and determining low-quality second inference results whose corresponding evaluation results meet the low-quality result conditions; Regarding both the high-quality second inference results and the low-quality second inference results as the reference inference results.
7. The method according to claim 6, characterized in that, Training the second language model based on the reference inference result includes: Obtaining the generation probability of the second inference path in the high-quality second inference result as the high-quality generation probability, and obtaining the generation probability of the second inference path in the low-quality second inference result as the low-quality generation probability; Training the second language model with the goal of maximizing the high-quality generation probability and minimizing the low-quality generation probability.
8. The method according to claim 7, wherein The evaluation result is the evaluation score determined by the evaluation model for the second inference result; training the second language model with the goal of maximizing the high-quality generation probability and minimizing the low-quality generation probability includes: Determining a positive loss according to the evaluation score of the high-quality second inference result and the generation probability of the second inference path in the high-quality second inference result; Determining a negative loss according to the evaluation score of the low-quality second inference result and the generation probability of the second inference path in the low-quality second inference result; Training the second language model based on the positive loss and the negative loss, where the positive loss is used to achieve the goal of maximizing the high-quality generation probability, and the negative loss is used to achieve the goal of minimizing the low-quality generation probability.
9. The method according to any one of claims 4 to 8, characterized in that Training the second language model based on the reference inference result to obtain a target language model includes: Performing iterative multi-round alternating training on the second language model and the evaluation model, and using the second language model obtained in the last round of the alternating training as the target language model; Wherein, in each round of the alternating training, the second language model is trained based on the reference inference result, and the evaluation model is trained based on multiple third inference results generated by the second language model.
10. The method according to claim 9, characterized in that, Training the evaluation model based on multiple third inference results generated by the second language model includes: According to the third input data, determining the multiple third inference results through the second language model, and determining the generation probabilities respectively corresponding to the multiple third inference results; Determining the annotation scores respectively corresponding to the multiple third inference results according to the generation probabilities respectively corresponding to the multiple third inference results; Determining the prediction scores respectively corresponding to the multiple third inference results through the evaluation model according to the multiple third inference results; Training the evaluation model according to the prediction scores and annotation scores respectively corresponding to the multiple third inference results.
11. The method according to any one of claims 1 to 10, characterized in that The method further includes: Among the multiple second inference results, determine a reasonable second inference result whose corresponding evaluation result meets the reasonable result condition, and update the reasonable inference result by using the reasonable second inference result; Training the second language model based on the reference inference result to obtain a target language model includes: Training the second language model based on the reference inference result to obtain a third language model; Training the third language model based on the updated reasonable inference result to obtain the target language model.
12. The method according to claim 11, wherein Training the third language model based on the updated reasonable inference result to obtain the target language model includes: Training the third language model based on the updated reasonable inference result to obtain a fourth language model; Determining multiple fourth inference results through the fourth language model according to fourth input data; Determining the evaluation result corresponding to each of the multiple fourth inference results; Selecting at least one of the multiple fourth inference results from the multiple fourth inference results according to the evaluation result corresponding to each of the multiple fourth inference results, and updating the reference inference result; Training the fourth language model based on the updated reference inference result to obtain the target language model.
13. A model training device, characterized in that, The device includes: A first inference module, configured to determine multiple first inference results through a first language model according to first input data, where the first inference results include a first inference path of the first language model for the first input data and a first inference conclusion generated by the first language model based on the first input data; A first screening module, configured to determine a reasonable inference result among the multiple first inference results according to a preset discrimination rule; A first training module, configured to train the first language model based on the reasonable inference result to obtain a second language model; A second inference module, configured to determine multiple second inference results through the second language model according to second input data, where the second inference results include a second inference path of the second language model for the second input data and a second inference conclusion generated by the second language model based on the second input data; An evaluation module, configured to determine the evaluation result corresponding to each of the multiple second inference results; A second screening module, configured to determine a reference inference result among the multiple second inference results according to the evaluation result corresponding to each of the multiple second inference results; A second training module, configured to train the second language model based on the reference inference result to obtain a target language model.
14. A computer device, characterized in that, The device includes a processor and a memory; The memory is used to store a computer program; The processor is configured to execute the model training method according to any one of claims 1 to 12 based on the computer program.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer program is executed by an electronic device, the model training method according to any one of claims 1 to 12 is implemented.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the model training method according to any one of claims 1 to 12 is implemented.
Citation Information
Cited By
Text classification model training method, system and equipment based on reinforcement learning
CN121636700A
Data processing method and device, electronic equipment, storage medium and product
CN122287917A