Data processing method and device, product and equipment
By directly using the target reference response result to guide the first language model to perform prediction processing when the language model prediction fails, the problems of low training efficiency and poor effect in the existing technology are solved, and more efficient and accurate training results are achieved.
Patent Information
- Application Number
- CN202510704429.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-09
AI Technical Summary
In the existing technology, language models may need multiple predictions during training to generate correct response results, resulting in poor training results and low efficiency.
The first language model is obtained and the first sample query data with label information is selected from the sample query dataset. The first language model is called for prediction processing. If the prediction fails, the target reference response result is directly used to guide the first language model for prediction processing to obtain the second language model.
This improves the efficiency and accuracy of training the second language model, avoids repeated prediction processes, and improves training results.
Smart Images

Figure CN120611792A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a data processing method, apparatus, product, and device. Background Art
[0002] With the continuous development of computer networks, artificial intelligence has been applied in more and more fields. For example, in the field of question answering, language models can be used to understand user questions and give users corresponding responses.
[0003] In existing applications, language models are trained by repeatedly predicting and generating responses to input data until a correct response is generated. These responses are then used to correct the model parameters. However, in these cases, the language model may need to predict the input data many times before generating a correct response, or may fail to generate a correct response at all. This results in poor training effectiveness and low training efficiency. Summary of the Invention
[0004] This application provides a data processing method, apparatus, product, and device that can improve the efficiency and accuracy of training a second language model.
[0005] On one hand, the present application provides a data processing method, the method comprising:
[0006] Obtaining a first language model, and selecting first sample query data for training the first language model from a sample query dataset, wherein the first sample query data has label information, and the label information is used to indicate a target reference response result corresponding to the first sample query data;
[0007] Calling the first language model to perform prediction processing on the first sample query data;
[0008] If the first language model fails to predict the first sample query data, the first language model is guided to perform prediction processing on the first sample query data based on the target reference response result to obtain a second language model, which is used to generate a corresponding prediction response result based on the input query data.
[0009] In one embodiment, the target reference response-based guidance of the first language model to perform prediction processing on the first sample query data to obtain the second language model includes:
[0010] inputting the target reference response result into the first language model;
[0011] The first language model is called to learn a prediction process of using the first sample query data to predict the target reference response result to obtain a second language model.
[0012] In one embodiment, the prediction process of the first language model learning to predict the target reference response result using the first sample query data is a process of inferring the correct prediction path for predicting the target reference response result using the first sample query data; and
[0013] The process of the first language model inferring the correct prediction path is a process of optimizing the model parameters of the first language model.
[0014] In one embodiment, calling the first language model to perform prediction processing on the first sample query data includes:
[0015] Calling the first language model to perform prediction processing on the first sample query data to generate N first prediction response results for the first sample query data;
[0016] If the N first predicted response results do not include the target reference response result, it is determined that the first language model fails to predict the first sample query data;
[0017] If the N first predicted response results include the target reference response result, it is determined that the first language model has successfully predicted the first sample query data.
[0018] In one aspect, the present application provides a data processing device, comprising:
[0019] an acquisition module, configured to acquire a first language model and select first sample query data from a sample query dataset for training the first language model, wherein the first sample query data has label information, and the label information is used to indicate a target reference response result corresponding to the first sample query data;
[0020] A prediction module, configured to call the first language model to perform prediction processing on the first sample query data;
[0021] The guidance module is used to guide the call of the first language model to perform prediction processing on the first sample query data based on the target reference response result if the first language model fails to predict the first sample query data, thereby obtaining a second language model. The second language model is used to generate a corresponding prediction response result based on the input query data.
[0022] In one embodiment, the first language model generates N first predicted response results after performing prediction processing on the first sample query data, and each first predicted response result has its own prediction probability, where N is a positive integer;
[0023] The guidance module guides the first language model to perform prediction processing on the first sample query data based on the target reference response result to obtain the second language model, including:
[0024] Determine the first predicted response result with the highest prediction probability among the N first predicted response results as the output response result;
[0025] Based on the output response result and the target reference response result, the first language model is guided to perform prediction processing on the first sample query data to obtain a second language model.
[0026] In one embodiment, the guidance module guides the first language model to perform prediction processing on the first sample query data based on the output response result and the target reference response result to obtain the second language model, including:
[0027] Searching for negative knowledge information associated with the first sample query data from a preset knowledge graph based on the output response result; and
[0028] Searching the knowledge graph for positive knowledge information associated with the first sample query data based on the target reference response result;
[0029] Negative knowledge information and positive knowledge information are used to guide the first language model to perform prediction processing on the first sample query data to obtain a second language model.
[0030] In one embodiment, the first sample query data is prompt information of the first language model; and the guidance module uses the negative knowledge information and the positive knowledge information to guide the first language model to perform prediction processing on the first sample query data to obtain the second language model, including:
[0031] Using the negative knowledge information and the positive knowledge information as background information of the first sample query data to construct reconstruction prompt information for the first language model, where the reconstruction prompt information includes the first sample query data and the background information;
[0032] Inputting the reconstructed prompt information into the first language model, and calling the first language model to perform prediction processing on the reconstructed prompt information to generate an updated prediction response result for the first sample query data;
[0033] Based on the difference between the updated predicted response result and the target reference response result, the model parameters of the first language model are modified to obtain the second language model.
[0034] In one embodiment, the knowledge graph is constructed based on multiple entities and attribute information of multiple entities;
[0035] The method of guiding the module to search for negative knowledge information associated with the first sample query data from a preset knowledge graph based on the output response result includes:
[0036] Obtain the first entity associated with the output reply result, and search the knowledge graph for attribute information of the first entity as negative knowledge information;
[0037] And, the method of guiding the module to search the knowledge graph for positive knowledge information associated with the first sample query data based on the target reference response result includes:
[0038] Obtain the second entity associated with the target reference reply result, and search the attribute information of the second entity from the knowledge graph as positive knowledge information.
[0039] In one embodiment, the guidance module guides the first language model to perform prediction processing on the first sample query data based on the target reference response result to obtain the second language model, including:
[0040] inputting the target reference response result into the first language model;
[0041] The first language model is called to learn a prediction process of using the first sample query data to predict the target reference response result to obtain a second language model.
[0042] In one embodiment, the prediction process of the first language model learning to predict the target reference response result using the first sample query data is a process of inferring the correct prediction path for predicting the target reference response result using the first sample query data; and
[0043] The process of the first language model inferring the correct prediction path is a process of optimizing the model parameters of the first language model.
[0044] In one embodiment, the prediction module calls the first language model to perform prediction processing on the first sample query data, including:
[0045] Calling the first language model to perform prediction processing on the first sample query data to generate N first prediction response results for the first sample query data;
[0046] If the N first predicted response results do not include the target reference response result, it is determined that the first language model fails to predict the first sample query data;
[0047] If the N first predicted response results include the target reference response result, it is determined that the first language model successfully predicts the first sample query data.
[0048] In one embodiment, the method in which the acquisition module selects first sample query data for training the first language model from the sample query data set includes:
[0049] Performing difficulty classification on the sample query data in the sample query data set to obtain sample query data of M difficulty levels, where M is a positive integer;
[0050] Selecting first sample query data based on the sample query data of M difficulty levels;
[0051] The M difficulty levels include the first difficulty level to the Mth difficulty level, and the first difficulty level to the Mth difficulty level increase in sequence.
[0052] In one embodiment, the first language model is an initial language model or a language model obtained by training the initial language model;
[0053] The acquisition module performs difficulty classification processing on the sample query data in the sample query data set to obtain the sample query data of M difficulty levels, including:
[0054] Calling the initial language model to perform prediction processing on each sample query data in the sample query data set to generate a prediction index value for each sample query data;
[0055] The sample query data in the sample query data set are divided into difficulty levels based on the prediction index value of each sample query data to obtain sample query data of M difficulty levels.
[0056] In one embodiment, any one of the sample query data sets is second sample query data, and the prediction indicator value of the second sample query data includes at least one of the following:
[0057] The prediction loss value of the initial language model for the second sample query data, and the prediction entropy value of the initial language model for the second sample query data;
[0058] The predicted entropy value refers to the entropy value of the predicted probability distribution generated by the initial language model for the second sample query data, and the predicted probability distribution is composed of the predicted probabilities of N predicted response results generated by the initial language model for the second sample query data.
[0059] In one embodiment, if the prediction index value of the second sample query data is a predicted loss value, the acquisition module performs difficulty classification processing on the sample query data in the sample query data set based on the prediction index value of each sample query data to obtain sample query data of M difficulty levels, including:
[0060] Obtain loss value division ranges corresponding to M difficulty levels, where the M loss value division ranges corresponding to the M difficulty levels are continuous and non-overlapping;
[0061] Obtain a target loss value partition range in which the predicted loss value of the second sample query data lies among the M loss value partition ranges;
[0062] The difficulty level corresponding to the target loss value division range is determined as the difficulty level of the second sample query data.
[0063] In one embodiment, if the prediction index value of the second sample query data is a predicted entropy value, the acquisition module performs difficulty classification processing on the sample query data in the sample query data set based on the prediction index value of each sample query data to obtain sample query data of M difficulty levels, including:
[0064] Obtaining entropy value division ranges corresponding to M difficulty levels, wherein the M entropy value division ranges corresponding to the M difficulty levels are continuous and non-overlapping;
[0065] Obtain a target entropy value partition range in which the predicted entropy value of the second sample query data lies in the M entropy value partition ranges;
[0066] The difficulty level corresponding to the target entropy value division range is determined as the difficulty level of the second sample query data.
[0067] In one embodiment, if the prediction index value of the second sample query data includes a prediction loss value and a prediction entropy value, the acquisition module performs difficulty classification processing on the sample query data in the sample query data set based on the prediction index value of each sample query data to obtain sample query data of M difficulty levels, including:
[0068] Obtain index value division ranges corresponding to M difficulty levels, wherein the M index value division ranges corresponding to the M difficulty levels are continuous and non-overlapping;
[0069] Calculating the sum of the predicted loss value and the predicted entropy value of the second sample query data, and determining the calculated sum as the difficulty classification value of the second sample query data;
[0070] Obtain a target indicator value division range in which the difficulty division value of the second sample query data lies in the M indicator value division ranges;
[0071] The difficulty level corresponding to the target indicator value division range is determined as the difficulty level of the second sample query data.
[0072] In one embodiment, the first language model is an initial language model or a language model obtained by training the initial language model. When the initial language model is initially trained, sample query data of the first difficulty level among the sample query data of the M difficulty levels is used for training.
[0073] In the training process of the initial language model, the difficulty level of the sample query data used in each round of iterative training of the initial language model is dynamically changed according to the change in the prediction accuracy of the sample query data of the initial language model in each round of iterative training;
[0074] The first language model is obtained after the initial language model is trained for the Kth round of iterations, where K is a non-negative integer.
[0075] In one embodiment, the acquisition module selects the first sample query data based on the sample query data of M difficulty levels, including:
[0076] Obtain the i-th difficulty level of the sample query data used in the K-th round of iterative training, where i is a positive integer and is less than or equal to M;
[0077] Obtain the target prediction accuracy of the initial language model in training for the sample query data during the Kth round of iterative training;
[0078] According to the target prediction accuracy and the i-th difficulty level, the first sample query data is selected from the sample query data of M difficulty levels.
[0079] In one embodiment, the initial language model in training predicts N predicted response results for each of the H sample query data in the Kth round of iterative training, and each of the H sample query data has its own reference response result, where H is a positive integer;
[0080] The acquisition module obtains the target prediction accuracy of the initial language model in training for the sample query data during the Kth round of iterative training, including:
[0081] Count the number of reference response results in each of the N predicted response results for H sample query data, and the number of H corresponding reference response results for each of the H sample query data;
[0082] Calculate the average number among H numbers as the target prediction accuracy.
[0083] In one embodiment, the acquisition module selects the first sample query data from the sample query data of M difficulty levels according to the target prediction accuracy and the i-th difficulty level, including:
[0084] If the target prediction accuracy meets the training effect improvement indication range corresponding to the initial language model, and i is not equal to M, then the sample query data of the i+1th difficulty level is selected from the sample query data of the M difficulty levels as the first sample query data;
[0085] If the target prediction accuracy meets the training effect improvement indication range, and i is equal to M, then the sample query data of the i-th difficulty level is selected from the sample query data of the M difficulty levels as the first sample query data;
[0086] If the target prediction accuracy meets the training effect stagnation indication range corresponding to the initial language model, then select the sample query data of the i-th difficulty level from the sample query data of the M difficulty levels as the first sample query data;
[0087] If the target prediction accuracy meets the training effect degradation indication range corresponding to the initial language model, and i is not equal to 1, then the sample query data of the i-1th difficulty level is selected from the sample query data of the M difficulty levels as the first sample query data;
[0088] If the target prediction accuracy meets the training effect degradation indication range, and i is equal to 1, sample query data of the i-th difficulty level is selected from the sample query data of the M difficulty levels as the first sample query data.
[0089] In one embodiment, the guidance module is further configured to:
[0090] Obtaining query data sent by the client and inputting the query data into the second language model;
[0091] Calling the second language model to perform prediction processing on the input query data, generating N second predicted response results for the query data and a prediction probability of each second predicted response result;
[0092] Determine the second predicted response result with the highest prediction probability among the N second predicted response results as the target predicted response result for the query data;
[0093] The target prediction response result is returned to the client, so that the client outputs the target prediction response result.
[0094] In one aspect, the present application provides a computer device including a memory and a processor, wherein the memory stores a computer program. When the computer program is executed by the processor, the processor executes the method in one aspect of the present application.
[0095] In one aspect, the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the method in the above aspect.
[0096] In one aspect, the present application provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in various optional embodiments of the above aspect.
[0097] The present application can obtain a first language model and select first sample query data for training the first language model from a sample query data set, the first sample query data having label information, and the label information being used to indicate a target reference response result corresponding to the first sample query data; calling the first language model to perform prediction processing on the first sample query data; and, if the first language model fails to predict the first sample query data, guiding the first language model to learn to perform prediction processing on the first sample query data based on the target reference response result, thereby obtaining a second language model, which can be used to generate a corresponding prediction response result based on the input query data. It can be seen that the method proposed in this application can directly guide the first language model to predict the first sample query data when the first language model fails to predict the first sample query data through the correct answer to the prediction processing of the first sample query data (that is, the target reference response result corresponding to the first sample query data), so that the first language model can quickly learn the process of correctly predicting the first sample query data from the correct direction, thereby obtaining a trained second language model, so that the first language model does not need to repeatedly predict the first sample query data multiple times, thereby improving the efficiency of training the second language model. Moreover, since the correct answer directly guides the first language model to learn the process of correctly predicting the first sample query data, the accuracy of the trained second language model is also guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0099] Figure 1 This is a schematic diagram of the network architecture of a question-answering network provided in an embodiment of the present application;
[0100] Figure 2 This is a schematic diagram of a scenario for training a second language model provided in an embodiment of the present application;
[0101] Figure 3 This is a flow chart of a data processing method provided in an embodiment of the present application;
[0102] Figure 4 This is a schematic diagram of a scenario for predicting and processing first sample query data provided by an embodiment of the present application;
[0103] Figure 5This is a schematic diagram of a process for selecting first sample query data from a sample query data set provided by an embodiment of the present application;
[0104] Figure 6 This is a schematic diagram of a scenario for classifying the difficulty level of second sample query data provided by an embodiment of the present application;
[0105] Figure 7 This is a schematic diagram of a scenario of selecting first sample query data from a sample query data set provided by an embodiment of the present application;
[0106] Figure 8 This is a flow chart of a model training process provided by an embodiment of the present application;
[0107] Figure 9 is a structural diagram of a data processing device provided in an embodiment of the present application;
[0108] Figure 10 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0109] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0110] All data collected in this application (such as language models, sample query data, query data, and predicted response results and other related data) are collected with the consent and authorization of the owner of the data (such as users, institutions or enterprises), and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of the relevant regions.
[0111] Here, the relevant technical concepts involved in this application are explained:
[0112] GRPO: Group Relative Policy Optimization, a reinforcement learning algorithm for models.
[0113] Course Learning Optimization Strategy: This application provides an innovative strategy that can divide training data into different difficulty levels and gradually train the model according to the training data of each difficulty level, so that the model can gradually adapt to the learning of more complex training data, improve the efficiency of model training, and make it easier for the model to converge.
[0114] Reasoning process optimization based on answer feedback: This application provides an innovative method that directly provides the correct answer to the model when the model fails to predict the correct answer (i.e., a prediction failure). By allowing the model to deduce each step of the reasoning process itself, the model is forced to correct the error and generate the correct reasoning path. In this way, the model can master the correct reasoning method for the training data at the initial stage, avoiding the inefficient process of repeating the prediction when the prediction fails.
[0115] Large Language Model (LLM). This is an AI model built on deep neural networks (specifically the Transformer architecture). It learns language patterns through pre-training on massive amounts of text data and can perform tasks such as natural language understanding (NLU), generation (NLG), and reasoning.
[0116] Entity: It can refer to things or objects that exist objectively and can be distinguished from each other. It can represent actual things (such as goods, students) or abstract concepts (such as courses, competitions).
[0117] Knowledge graph: A technical framework that describes real-world entities and their relationships in the form of a structured semantic network to build a relational knowledge base and support efficient semantic reasoning and knowledge retrieval. The data in the knowledge graph is organized as a graph (consisting of nodes and edges) rather than a traditional table, thus supporting complex relational queries.
[0118] Rag: A framework that enhances model generation capabilities by retrieving external knowledge bases. It aims to address the knowledge limitations, hallucination suppression, and context expansion issues of large language models (such as LLMs). RAG can significantly improve the practicality and reliability of LLMs in professional scenarios.
[0119] See Figure 1 , Figure 1 This is a schematic diagram of the network architecture of a question-answering network provided in an embodiment of the present application. Figure 1 As shown, the network architecture may include a server 200 and a terminal device cluster, and the terminal device cluster may include one or more terminal devices, and the number of terminal devices is not limited here. Figure 1 As shown, the multiple terminal devices may specifically include terminal device 1, terminal device 2, terminal device 3, ..., terminal device n; Figure 1 As shown, terminal device 1, terminal device 2, terminal device 3, ..., terminal device n can all be connected to the server 200 through a network, so that each terminal device can exchange data with the server 200 through the network connection.
[0120] like Figure 1The server 200 shown can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (content distribution network), and big data and artificial intelligence platforms. The terminal device can be: a smart phone, tablet computer, laptop computer, desktop computer, smart TV, car terminal, smart home terminal, etc. The following takes the communication between the terminal device 1 and the server 200 as an example to describe the embodiment of the present application in detail.
[0121] The terminal device 1 may have a client, such as a question-and-answer client, and the server 200 may be the background device of the client. The client can obtain the inquiry data input by the user (such as the inquiry data may be a question), and the client can give the inquiry data to the server 200 to request the server 200 to generate a predicted reply result corresponding to the inquiry data (such as the predicted reply result may be an answer to the question). The predicted reply result corresponding to the inquiry data may be predicted and generated by the server 200 by calling a trained second language model for the inquiry data. The server 200 can return the generated predicted reply result to the terminal device 1, that is, return it to the client in the terminal device 1, so that the client can output the predicted reply result in the client interface for the user to view.
[0122] Please also see Figure 2 , Figure 2 This is a schematic diagram of a scenario for training a second language model provided in an embodiment of the present application. Figure 2 The above-mentioned Figure 1 The server 200 in the process of training to obtain the second language model. Figure 2 As shown, the server 200 can obtain a first language model that needs to be trained, and can select sample query data from the sample query data set for training the first language model. The selected sample query data can be referred to as first sample query data.
[0123] Server 200 may use the first sample query data to train a first language model, for example, by calling the first language model to perform prediction processing on the first sample query data. The first sample query data may have a corresponding target reference response result, which may be an ideal response result that is desired to be predicted for the first sample query data. The target reference response result may be understood as the correct answer to the prediction processing of the first sample query data.
[0124] If the server 200 fails to predict the first sample query data, the server 200 can directly use the target reference response result to guide the first language model, that is, to guide the first language model to perform correct prediction processing on the first sample query data, thereby training the second language model. In other words, in this application, when the first language model predicts the first sample query data, it only needs to perform a one-time prediction of multiple predicted response results. If the prediction fails, there is no need to predict the predicted response results again.
[0125] Through the above-mentioned method of the present application, when the first language model fails to predict the first sample query data, the target reference response result corresponding to the first sample query data can be directly given to the first language model, so that the first language model can quickly and correctly learn the correct prediction path for predicting the first sample query data, thereby improving the efficiency of training the second language model and ensuring the accuracy of training the second language model.
[0126] See Figure 3 , Figure 3 This is a flow chart of a data processing method provided by an embodiment of the present application. The execution subject in the embodiment of the present application can be a data processing device (which can be referred to as a processing device), and the processing device can be a computer device or a computer device cluster composed of multiple computer devices. The computer device can be a server, a terminal device, or other devices, without limitation. Figure 3 As shown, the method may include:
[0127] Step S101: obtain a first language model and select first sample query data for training the first language model from a sample query data set. The first sample query data has label information, and the label information is used to indicate a target reference response result corresponding to the first sample query data.
[0128] In one embodiment, a processing device may obtain a first language model, which is a language model to be trained. The first language model may be an initial language model or a language model obtained by training the initial language model (but not yet fully trained). The initial language model is the very first language model to be trained. That is, the purpose of this application may be to train the initial language model to obtain a finally trained second language model, and the first language model may be a language model that has not yet been fully trained at some point during the training of the initial language model and before the second language model is obtained.
[0129] The initial language model can be a pre-trained language model (e.g., a language model pre-trained with some general world knowledge), that is, the initial language model can be a pre-trained language model that needs to be further trained (or trained in a specific training direction). For example, the initial language model can be a language model that has been trained on the market, but when applied to an actual business scenario, it needs to be further strengthened with the business data related to the business scenario. The initial language model can be a large language model.
[0130] The sample query dataset can contain a large amount of sample data used to train the initial language model. This sample data can be sample query data, that is, the sample query dataset can contain a large amount of sample query data. The sample query data can be data used to input into the language model to obtain a corresponding response result. For example, the sample query data can be a text question, a question that combines text and images, or a question about video data, etc.
[0131] For example, a sample query data could be the question "How to improve learning efficiency?" or it could include an image (can be any color image) and a question about that image, such as "How many colors are there in this image?" The specific nature of the sample query data can be determined based on the actual requirements for language model training.
[0132] The processing device can select sample query data for training the first language model from the sample query data set, and the selected sample query data can be referred to as the first sample query data. There can be multiple first sample query data, and the first sample query data can have label information, and the label information is used to indicate the reference response result corresponding to the first sample query data. The reference response result of the first sample query data can be referred to as the target reference response result. The target reference response result can be the ideal response result that is desired to be obtained by performing predictive processing on the first sample query data, that is, the target reference response result can be understood as the correct answer (that is, the correct output) for the predictive processing of the first sample query data.
[0133] Specifically, how to select the first sample query data for training the first language model from the sample query data set can be found in the following Figure 5 Description of the content in the corresponding embodiment: Since the logic of predicting and processing each first sample query data is the same, the following is a specific description of the process of predicting and processing one first sample query data as an example.
[0134] Step S102: calling a first language model to perform prediction processing on the first sample query data.
[0135] In one embodiment, the processing device can call the first language model to perform prediction processing on the first sample query data selected above. The process may include: the processing device can call the first language model to perform prediction processing on the first sample query data, and can generate N prediction response results for the first sample query data. The N prediction response results can be referred to as N first prediction response results, where N is a positive integer. The specific value of N can be set according to actual needs. N is the number of prediction response results that the first language model can predict for one sample query data at a time.
[0136] If the N first predicted response results do not include the target reference response result, it can be determined that the first language model failed to predict the first sample query data, and therefore the first language model did not predict the correct answer to the first sample query data.
[0137] If the N first predicted response results include the target reference response result, it can be determined that the first language model successfully predicted the first sample query data, and therefore the first language model predicted the correct answer to the first sample query data.
[0138] Furthermore, if the first language model successfully predicts the first sample query data, the processing device can directly correct the model parameters of the first language model by using the first predicted response result that is consistent with the target reference response result among the N first predicted response results. This process can be to use the original loss function of the first language model to correct the model parameters of the first language model. For example, each of the N predicted first predicted response results can have a prediction probability (i.e., prediction confidence). Therefore, by correcting the model parameters of the first language model, the prediction probability of the first predicted response result that is consistent with the target reference response result among the N first predicted response results can be higher, such as approaching the highest value of 1, thereby achieving the correction of the model parameters of the first language model.
[0139] See Figure 4 , Figure 4 This is a schematic diagram of a scenario for predicting and processing the first sample query data provided by an embodiment of the present application. Figure 4 As shown, the first sample query data can be input into the first language model to call the first language model to perform predictive processing on the first sample query data, and N first predicted reply results can be generated for the first sample query data, where N can be equal to 8, and the N first predicted reply results can include first predicted reply results g1 to first predicted reply results g8.
[0140] Therefore, if the first predicted response result g1 to the first predicted response result g8 includes the target reference response result corresponding to the first sample inquiry data, that is, there is a first predicted response result that is consistent with the target reference response result corresponding to the first sample inquiry data, then it can be considered that the prediction of the first sample inquiry data is successful.
[0141] However, if the first predicted response results g1 to g8 do not include the target reference response result corresponding to the first sample inquiry data, that is, there is no first predicted response result consistent with the target reference response result corresponding to the first sample inquiry data, then it can be considered that the prediction of the first sample inquiry data has failed.
[0142] Step S103: If the first language model fails to predict the first sample query data, the first language model is guided to perform prediction processing on the first sample query data based on the target reference response result to obtain a second language model, which is used to generate a corresponding prediction response result based on the input query data.
[0143] In one embodiment, if the first language model fails to predict the first sample query data, the processing device can directly use the correct answer (i.e., the target reference response result) to guide the first language model to predict the first sample query data, thereby obtaining a trained second language model. The embodiments of this application exemplarily provide the following method to guide the first language model to predict the first sample query data using the target reference response result, as described below.
[0144] For example, if the first language model fails to predict the first sample query data, the processing device can directly give the correct answer (ie, the target reference response result) to the first language model, that is, the target reference response result can be input into the first language model.
[0145] The processing device can call upon the first language model to learn a prediction process for predicting the input target reference response using the first sample query data. The process of learning this prediction process is the process of guiding (i.e., directing) the first language model to perform predictive processing on the first sample query data based on the given target reference response. During this process, the first language model can continuously modify (i.e., optimize) its own model parameters so that the target reference response can be predicted using the first sample query data. In other words, the process of the first language model learning to predict the target reference response using the first sample query data is the process of modifying the model parameters of the first language model.
[0146] Among them, the prediction process of the first language model learning to use the first sample inquiry data to predict the target reference response result is also the process of inferring the correct prediction path for predicting the target reference response result using the first sample inquiry data. The correct prediction path is the path that the large language model can predict the target reference response result using the first sample inquiry data. The correct prediction path can be understood as the thinking process of the first language model using the first sample inquiry data to predict the target reference response result. That is, when the first language model fails to predict the first sample inquiry data, the present application can guide the first language model to reason and predict the first sample inquiry data from the correct reasoning direction by directly giving the target reference response result corresponding to the first sample inquiry data to the first language model, that is, allowing the first language model to learn the correct thinking and process for the first sample inquiry data, thereby predicting the correct predicted response result corresponding to the first sample inquiry data (that is, the predicted response result consistent with the target reference response result).
[0147] The process of the first language model inferring the correct prediction path is the process of optimizing (i.e., correcting) the model parameters of the first language model. That is, during the process of the first language model inferring the correct prediction path, the model parameters of the first language model can be continuously corrected. By directly providing the correct answer (i.e., the target reference response result) to the first language model as described above, instructing the first language model to predict the first sample query data in the direction of the correct answer, the efficiency of the first language model in correctly learning the first sample query data can be greatly improved, thereby greatly improving the efficiency of obtaining the second language model through training with the first language model, ensuring the efficiency and reliability of model training.
[0148] For example, in another embodiment, if the first language model fails to predict the first sample query data, the processing device may not directly give the target reference response result to the first language model, but may use the target reference response result to guide the first language model to generate a new predicted response result for the first sample query data, thereby correcting the model parameters of the first language model through the new predicted response result to obtain a second language model, as described below.
[0149] As described in step S102 above, after the first language model performs predictive processing on the first sample query data, it can generate N first predicted response results, each of which can have a respective prediction probability. Therefore, the processing device can use the first predicted response result with the highest prediction probability among the N first predicted response results as the output response result. This output response result can be an incorrect response result generated by the first language model for the first sample query data.
[0150] The processing device may use the output response result and the target reference response result to guide the first language model to perform prediction processing on the first sample query data to obtain a second language model. This process may include:
[0151] A knowledge graph may be preset in this application, and the knowledge graph may be used to assist in training the initial language model. The knowledge graph may be composed of multiple entities and attribute information of the multiple entities. For example, the knowledge graph may include entity nodes for representing the multiple entities and information nodes for the attribute information of the multiple entities. An entity node and an information node for its attribute information may have an edge (i.e., a connection relationship). Among them, an entity may have one or more attribute information, and an attribute information of an entity may have an information node, i.e., the attribute information of an entity may be represented by one or more information nodes in the knowledge graph.
[0152] Entities can refer to objectively existing and distinguishable things or objects, representing both actual things (such as commodities and students) and abstract concepts (such as courses and competitions). The specific entities and attribute information included in the above knowledge graph can be customized based on the actual needs of auxiliary model training, and this application does not impose any restrictions on this.
[0153] Therefore, the processing device can search the preset knowledge graph for negative knowledge information associated with the first sample query data using the output response result. That is, because the output response result is an incorrect response result predicted and generated by the first language model for the first sample query data, the relevant knowledge information searched for in the knowledge graph using the output response result can be regarded as negative knowledge information associated with the first sample query data (or can also be referred to as incorrect knowledge information), and subsequently used to negatively guide the first language model's prediction processing of the first sample query data.
[0154] If the processing device can obtain the entity associated with the output reply result, the entity associated with the output reply result can be referred to as the first entity, and the first entity can be one or more. The first entity can be an entity extracted from the output reply result, or the first entity can also be a related entity generated by summarizing the output reply result after the current first language model understands the output reply result. Therefore, the processing device can search for the attribute information of the first entity in the knowledge graph as the above-mentioned negative knowledge information, such as the negative knowledge information can be the attribute information represented by the information node in the knowledge graph that has a connection relationship with the entity node of the first entity.
[0155] Furthermore, the processing device may also search the knowledge graph for positive knowledge information associated with the first sample query data using the target reference response result. That is, because the target reference response result is the correct answer (i.e., the reference answer) for predicting the first sample query data, the relevant knowledge information searched for in the knowledge graph using the target reference response result may be considered as positive knowledge information associated with the first sample query data (or correct knowledge information), and subsequently used to provide positive guidance to the first language model's predictive processing of the first sample query data.
[0156] Similarly, the processing device can obtain the entity associated with the target reference reply result, and the entity associated with the target reference reply result can be called a second entity. The second entity can be one or more. The second entity can also be an entity extracted from the target reference reply result, or the second entity can also be a related entity generated by summarizing the target reference reply result after the current first language model understands the target reference reply result. Therefore, the processing device can search for the attribute information of the second entity in the knowledge graph as the above-mentioned positive knowledge information, such as the positive knowledge information can be the attribute information represented by the information node in the knowledge graph that has a connection relationship with the entity node of the second entity.
[0157] After obtaining the negative knowledge information and positive knowledge information associated with the first sample query data, the processing device can use the negative knowledge information and positive knowledge information to guide the first language model to perform predictive processing on the first sample query data, thereby obtaining a second language model, as described below.
[0158] The sample query data in this application may be prompt information (ie, prompt word) belonging to the first language model, that is, the first sample query data also belongs to the prompt information of the first language model.
[0159] The processing device may use the negative knowledge information and positive knowledge information as background information for the first sample query data to form reconstructed prompt information for the first language model. Specifically, the reconstructed prompt information is prompt information reconstructed from the first sample query data and may include the first sample query data and the background information (i.e., the negative knowledge information and the positive knowledge information).
[0160] In other words, reconstruction prompt information for the first language model can be constructed using the negative knowledge information, positive knowledge information, and the first sample query data. In this reconstruction prompt information, the negative knowledge information and positive knowledge information serve as background information for the first sample query data. This reconstruction prompt information can prompt the first language model to treat the negative knowledge information as erroneous information for predicting the first sample query data, and to treat the positive knowledge information as correct information for predicting the first sample query data, thereby guiding the first language model to predict the first sample query data in the correct direction.
[0161] Therefore, the reconstructed prompt information can be input into the first language model to invoke the first language model to perform prediction processing on the reconstructed prompt information. That is, prediction processing is performed on the first sample query data based on the background information in the reconstructed prompt information, thereby generating an updated predicted response result for the first sample query data. The updated predicted response result is the predicted response result newly predicted by the first language model for the first sample query data under the guidance of the above-mentioned negative knowledge information and positive knowledge information.
[0162] For example, the first language model can generate a reverse question for itself based on the background information (i.e., negative knowledge information and positive knowledge information). The reverse prompt can be a comparison question about the negative knowledge information and the positive knowledge information. Thus, the first language model can understand and summarize the key information used for predictive processing of the first sample query data by answering the reverse question itself.
[0163] For example, the output response result may be a cold, the target reference response result may be influenza, the negative knowledge information may be symptoms related to a cold (such as nasal congestion, runny nose, sore throat, etc.), and the positive knowledge information may be symptoms related to influenza (such as high fever, general fatigue, severe headache, etc.). Therefore, the reverse question generated by the negative knowledge information and the positive knowledge information may be "Did I overlook the key symptom of general fatigue related to influenza in the input data?" The input data may be the first sample query data. The first language model can obtain key information for predicting the first sample query data, such as general fatigue, by answering the reverse prompt. This reverse question is only an exemplary description. In actual application scenarios, the large language model can generate more diverse and accurate reverse questions based on the above background information based on the knowledge learned by itself, so as to achieve auxiliary understanding, summary and induction of the first sample query data through the background information, thereby obtaining key information for predicting the first sample query data.
[0164] The first language model can use the key information obtained based on the reverse question to correct the original prediction path for predicting the first sample query data, thereby achieving new prediction processing for the first sample query data along the correct prediction path to predict and generate the above-mentioned updated prediction response result for the first sample query data.
[0165] After predicting and generating the above-mentioned updated predicted reply result, the processing device can correct the model parameters of the first language model by the difference between the updated predicted reply result and the above-mentioned target reference reply result to obtain a second language model. For example, the initial language model can have its own loss function, and the processing device can substitute the updated predicted reply result and the target reference reply result into the loss function to obtain a predicted loss value for the updated predicted reply result. Subsequently, the model parameters of the first language model can be corrected by the predicted loss value to obtain a second language model. For example, the goal of correcting the model parameters of the first language model by the predicted loss value is to correct the model parameters of the first language model so that the predicted loss value can be minimized (such as approaching 0). The predicted loss value can be used to reflect the difference between the updated predicted reply result and the target reference reply result. The larger the predicted loss value, the greater the difference between the updated predicted reply result and the target reference reply result, and the less accurate the updated predicted reply result generated by the first language model; conversely, the smaller the predicted loss value, the smaller the difference between the updated predicted reply result and the target reference reply result, and the more accurate the updated predicted reply result generated by the first language model.
[0166] The principle of substituting the updated prediction response result and the target reference response result into the loss function is the same as the following Figure 5 In the corresponding embodiment, the principle of substituting the predicted response result finally generated for the second sample query data and the reference response result corresponding to the second sample query data into the loss function is the same.
[0167] The above-mentioned process of assisting the training of the initial language model by introducing the knowledge graph in the present application can be a process of applying the Rag operation to the initial language model. Through the Rag operation, the initial language model can learn the knowledge in the knowledge graph during the training process, thereby achieving a wider expansion of the learned knowledge. Through the above process, the auxiliary training of the first language model through the knowledge graph is realized, that is, the first language model can perform prediction processing on the first sample query data in the correct direction through the knowledge graph. As a result, during the training process of the initial language model, not only the knowledge related to the sample query data can be learned, but also the knowledge in the knowledge graph can be learned. Therefore, the breadth, richness and accuracy of the knowledge learned by the initial language model are improved, thereby improving the effect of training the initial language model.
[0168] In one embodiment, multiple first sample query data may be selected. For first sample query data that fails prediction, the reference response corresponding to the first sample query data can be used to guide the first language model's prediction processing of the first sample query data, thereby correcting the model parameters. For first sample query data that succeeds prediction, the model parameters can be directly corrected using the predicted response results for the first sample query data that are consistent with the reference response results corresponding to the first sample query data. This ensures the efficiency and accuracy of model training.
[0169] The processing device can continuously perform multiple rounds of iterative training on the first language model in the manner described above. When the training stop condition is reached, the second language model can be obtained. This second language model is the final trained (i.e., trained) language model. That is, the second language model is the language model ultimately trained with the initial language model. Subsequently, this second language model can be applied to actual data prediction services.
[0170] The training stop condition can be flexibly set according to actual business needs. For example, the training stop condition can be that the number of rounds of iterative training of the initial language model is equal to the set round threshold, that is, when the total number of rounds of iterative training of the initial language model is equal to the round threshold, the model training can be stopped, and the trained initial language model at this time can be used as the trained second language model; for example, the training stop condition can be that the precision and recall rate of the data prediction of the trained initial language model (including accuracy and recall rate) is higher than the set precision and recall threshold (including accuracy threshold and recall rate threshold). After each round of iterative training of the initial language model, the trained initial language model can be tested and obtained. The model performs data prediction (such as prediction processing of test query data, where the concept of the test query data has the same probability as the sample query data) with a precision recall rate. When the precision recall rate is greater than or equal to the precision recall threshold, model training can be stopped, and the trained initial language model at this time can be used as the trained second language model. For example, the training stopping condition can be training the model parameters of the initial language model to a convergence state, that is, when the initial language model is iteratively trained until the model parameters reach a convergence state, model training can be stopped, and the trained initial language model at this time can be used as the trained second language model; and so on.
[0171] It should be understood that the initial language model, first language model, and second language model mentioned above are all the same model. The initial language model is the model to be trained, which has not yet been trained using the method provided by this application. The first language model is the initial language model in training, and the second language model is the model after the initial language model is trained. In other words, the initial language model, the first language model, and the second language model are the same model, just corresponding to different stages of model training.
[0172] Furthermore, the processing device can obtain query data sent by the client. The query data can be data sent by the client for prediction, and the processing device can be the backend device of the client. The client can be any client that supports data prediction triggering, such as a webpage, application, or mini-program, or any other type of client, such as a question-and-answer client or a social client.
[0173] The concept and format of the query data are the same as those of the sample query data described above, such as a question that needs to be predicted, such as a question entered by a user in the client that needs a predicted reply result. The processing device can input the query data into the trained second language model to call the second language model to perform predictive processing on the input query data, and can generate N predicted reply results for the query data and a predicted probability (i.e., prediction confidence) for each predicted reply result. Among them, the N predicted reply results for the query data can be referred to as N second predicted reply results, and each second predicted reply result has its own corresponding prediction probability. The higher the predicted probability of the second predicted reply result, the higher the confidence of the second language model in predicting the second predicted reply result, and the more accurate the predicted second predicted reply result; conversely, the lower the predicted probability of the second predicted reply result, the lower the confidence of the second language model in predicting the second predicted reply result, and the less accurate the predicted second predicted reply result.
[0174] Therefore, the processing device can use the second predicted response result with the highest prediction probability among the N second predicted response results as the target predicted response result for the above-mentioned query data. The target predicted response result is also the response result finally predicted for the query data.
[0175] The processing device can return the target prediction reply result to the client, so that the client can output (such as output in the client interface, or output through language broadcast, etc.) the target prediction reply result for the corresponding user to view.
[0176] When the model's initial prediction fails to obtain the correct answer, the present application can directly provide the correct answer to the model and allow the model to deduce the intermediate reasoning process (i.e., the process of inferring the correct answer through sample query data). This method enables the model to correct the reasoning path of the sample query data and quickly master the correct reasoning process, thereby effectively reducing the inefficiency of allowing the model to perform multiple redundant reasonings and reducing the computing resources consumed by multiple redundant reasonings, avoiding the problem of the model sticking to the wrong reasoning path and repetitive prediction errors, allowing the model to move along the optimal path (i.e., the correct reasoning path) to master the correct reasoning process in a shorter time, significantly shortening the training time of the model, and avoiding the model's inability to adapt to complex model training tasks under a fixed number of reasoning times (such as continuing to make the next prediction after a prediction fails, and so on, until the set fixed number of reasoning times is reached), resulting in the limitation of not being able to correctly deduce the sample query data, so that the trained model can handle more complex and challenging reasoning problems. The present application does not rely on the fixed number of reasoning times when training the model. The model only needs to perform a one-time model reasoning (i.e., prediction processing for sample query data).
[0177] The present application can obtain a first language model and select first sample query data for training the first language model from a sample query data set, the first sample query data having label information, and the label information being used to indicate a target reference response result corresponding to the first sample query data; calling the first language model to perform prediction processing on the first sample query data; and, if the first language model fails to predict the first sample query data, guiding the first language model to perform prediction processing on the first sample query data based on the target reference response result, thereby obtaining a second language model, which can be used to generate a corresponding prediction response result based on the input query data. It can be seen that the method proposed in this application can directly guide the first language model to predict the first sample query data when the first language model fails to predict the first sample query data through the correct answer to the prediction processing of the first sample query data (that is, the target reference response result corresponding to the first sample query data), so that the first language model can quickly learn the process of correctly predicting the first sample query data from the correct direction, thereby obtaining a trained second language model, so that the first language model does not need to repeatedly predict the first sample query data multiple times, thereby improving the efficiency of training the second language model. Moreover, since the correct answer directly guides the first language model to learn the process of correctly predicting the first sample query data, the accuracy of the trained second language model is also guaranteed.
[0178] See Figure 5 , Figure 5This is a flow chart of selecting first sample query data from a sample query data set provided by an embodiment of the present application. Figure 5 As shown, the process may include:
[0179] Step S201 : performing difficulty classification processing on the sample query data in the sample query data set to obtain sample query data of M difficulty levels, where M is a positive integer.
[0180] In one embodiment, the processing device can perform difficulty division processing on the sample query data in the sample query data set to obtain sample query data divided into M difficulty levels, where M is a positive integer, and the specific value of M can be flexibly set according to the actual application scenario. That is, the specific number of difficulty levels of sample query data that the sample query data in the sample query data set needs to be divided into can be determined based on actual business needs, and this application does not impose any restrictions on this. For example, M can be equal to 3, indicating that the sample query data in the sample query data set needs to be divided into sample query data of 3 difficulty levels, such as the 3 difficulty levels can represent the difficulty level of simple difficulty, the difficulty level of medium difficulty, and the difficulty level of high difficulty, respectively.
[0181] The M difficulty levels may include the 1st difficulty level to the Mth difficulty level in sequence, and the difficulty levels may increase in sequence from the 1st difficulty level to the Mth difficulty level. For example, any difficulty level among the M difficulty levels may be represented as the jth difficulty level, and sample query data of the jth difficulty level is more difficult (i.e., more difficult) than sample query data of the j-1th difficulty level, and sample query data of the jth difficulty level is simpler (i.e., less difficult) than sample query data of the j+1th difficulty level.
[0182] The following examples provide several ways to perform difficulty classification processing on the sample query data in the sample query data set. It should be noted that in actual application scenarios, other appropriate methods can also be used to perform difficulty classification processing on the sample query data in the sample query data set, and this application does not impose any restrictions on this.
[0183] The processing device may invoke the initial language model to perform prediction processing on each sample query data in the sample query data set, and may generate a prediction index value for each sample query data. The prediction index value may be used to indicate the accuracy of the initial language model's prediction for each sample query data. The lower the accuracy of the initial language model's prediction for a sample query data, the greater the difficulty of the initial language model in predicting the sample query data. Conversely, the higher the accuracy of the initial language model's prediction for a sample query data, the easier it is for the initial language model to predict the sample query data.
[0184] Therefore, the processing device can perform difficulty classification processing on the sample query data in the sample query data set based on the prediction index value predicted by the initial language model for each sample query data in the sample query data set to obtain the sample query data of the above-mentioned M difficulty levels, as described below.
[0185] Any sample query data in the sample query data set can be referred to as second sample query data. Since the principle for obtaining the prediction index value for each sample query data is the same, the following description specifically uses obtaining the prediction index value for the second sample query data as an example. It is understood that the processing device can obtain the prediction index value for each sample query data in the sample query data set according to the following principle.
[0186] The prediction index value of the second sample query data may include at least one of the following: a prediction loss value of the initial language model for the second sample query data, and a prediction entropy value of the initial language model for the second sample query data.
[0187] A method for obtaining a prediction loss value of the initial language model for the second sample query data may include: performing prediction processing on the second sample query data by the initial language model to generate N predicted response results for the second sample query data and a predicted probability for each predicted response result; and determining the predicted response result with the highest predicted probability among the N predicted response results as the predicted response result ultimately generated by the initial language model for the second sample query data. Furthermore, the initial language model may have a loss function, and the processing device may substitute the predicted response result ultimately generated by the initial language model for the second sample query data and a reference response result corresponding to the second sample query data into the loss function to obtain a prediction loss value of the initial language model for the second sample query data. A higher prediction loss value indicates a greater error in the initial language model's prediction processing of the second sample query data, a less accurate prediction, and a greater difficulty for the initial language model to predict the second sample query data. Conversely, a lower prediction loss value indicates a smaller error in the initial language model's prediction processing of the second sample query data, a more accurate prediction, and a less difficult prediction for the initial language model to predict the second sample query data. As shown in the following formula, the predicted loss value L(x) of the second sample query data can be:
[0188] L(x) = Loss(f(x), y x ) (1)
[0189] Where x represents the second sample query data, Loss represents the loss function of the initial language model, f(x) represents the predicted response result generated by the initial language model for the second sample query data, and yx Indicates the reference response result corresponding to the second sample query data.
[0190] For example, the initial language model can be trained using the GRPO algorithm. Therefore, the loss function of the initial language model can be the loss function defined by the GRPO algorithm, and the loss function can be composed of a policy gradient term and a KL divergence (an indicator used to measure the difference between two probability distributions) constraint term.
[0191] The method for obtaining the predicted entropy value of the initial language model for the second sample query data may include: the predicted entropy value may refer to the entropy value of the predicted probability distribution generated by the initial language model for the second sample query data (which may be simply referred to as entropy), and the predicted probability distribution may be a probability distribution composed of the N predicted probabilities of the N predicted response results predicted and generated by the initial language model for the second sample query data. That is, the predicted entropy value of the initial language model for the second sample query data can be calculated by the N predicted probabilities of the N predicted response results predicted and generated for the second sample query data. The predicted entropy value is the entropy value of the predicted probability distribution. The predicted entropy value can reflect the uncertainty of the initial language model's prediction processing of the second sample query data. The higher the predicted entropy value, the higher the uncertainty of the initial language model's prediction processing of the second sample query data; conversely, the lower the predicted entropy value, the lower the uncertainty of the initial language model's prediction processing of the second sample query data. As shown in the following formula, the predicted entropy value H(p(x)) of the initial language model for the second sample query data can be:
[0192]
[0193] Where p(x) represents the above predicted probability distribution, x represents the second sample query data, and y v represents the vth predicted response result generated by the initial language model for the second sample query data, and the value range of v can be [1, M]. p(y v / x) represents the predicted probability of the vth predicted response result generated by the initial language model for the second sample query data. log represents the logarithm.
[0194] Therefore, if the prediction index value of the second sample query data is the above-mentioned prediction loss value, the process of the processing device classifying the difficulty level of the sample query data in the sample query data set according to the prediction index value of each sample query data may include:
[0195] The processing device can obtain the loss value division ranges corresponding to the above-mentioned M difficulty levels. One difficulty level can correspond to one loss value division range. The M loss value division ranges corresponding to the M difficulty levels are continuous with each other (i.e., the values are continuous with each other) and do not overlap with each other (i.e., different loss value division ranges will not contain the same value). The loss value division ranges corresponding to the M difficulty levels can be adaptively set in advance according to the actual business scenario. For example, the overall range of the predicted loss value for predicting each sample query data according to the initial language model can be roughly evenly divided. The set M loss value division ranges should be able to cover the predicted loss value of each sample query data. For example, the overall range of the predicted loss values of the initial language model for predicting each sample query data is around 10 to 70, M is equal to 3, and the 3 difficulty levels can include the simple difficulty level, the medium difficulty level and the high difficulty level in turn. Then the M loss value division ranges corresponding to the M difficulty levels can include [0, 30), [30, 50), and [50, +∞) in turn, where +∞ represents positive infinity. The simple difficulty level can correspond to the loss value division range [0, 30), the medium difficulty level can correspond to the loss value division range [30, 50), and the high difficulty level can correspond to the loss value division range [50, +∞).
[0196] The processing device can obtain the target loss value division range within which the predicted loss value of the second sample query data falls within the M loss value division ranges, i.e., the predicted loss value of the second sample query data falls within the target loss value division range. Therefore, the processing device can use the difficulty level corresponding to the target loss value division range as the difficulty level of the second sample query data. For example, if the difficulty level corresponding to the target loss value division range is difficulty level 1, then the second sample query data can be classified as sample query data of difficulty level 1.
[0197] If the prediction index value of the second sample query data is the above-mentioned prediction entropy value, the processing device may classify the difficulty level of the sample query data in the sample query data set according to the prediction index value of each sample query data, which may include:
[0198] Similarly, the processing device can obtain the entropy value division ranges corresponding to the above-mentioned M difficulty levels, one difficulty level can correspond to one entropy value division range, and the M entropy value division ranges corresponding to the M difficulty levels are continuous with each other (i.e., the values are continuous with each other) and do not overlap with each other (i.e., different entropy value division ranges will not contain the same value). The entropy value division ranges corresponding to the M difficulty levels can be adaptively set in advance according to the actual business scenario. For example, the overall range of the predicted entropy value for predicting each sample query data according to the initial language model can be roughly evenly divided, and the set M entropy value division ranges should be able to cover the predicted entropy value of each sample query data. Similarly, for example, the overall range of the predicted entropy value of the initial language model for predicting each sample query data is around 0 to 60, M is equal to 3, and the 3 difficulty levels can include the simple difficulty level, the medium difficulty level and the high difficulty level in turn. Then the M entropy value division ranges corresponding to the M difficulty levels can include [0, 20), [20, 40), and [40, +∞) in turn. The simple difficulty level can correspond to the loss value division range [0, 20), the medium difficulty level can correspond to the loss value division range [30, 40), and the high difficulty level can correspond to the loss value division range [40, +∞).
[0199] The processing device can obtain the target entropy value division range in which the predicted entropy value of the second sample query data falls within the M entropy value division ranges, that is, the predicted entropy value of the second sample query data falls within the target entropy value division range. Therefore, the processing device can use the difficulty level corresponding to the target entropy value division range as the difficulty level of the second sample query data. For example, if the difficulty level corresponding to the target entropy value division range is the third difficulty level, then the second sample query data can be classified as sample query data of the third difficulty level.
[0200] Furthermore, if the prediction index value of the second sample query data includes the above-mentioned prediction loss value and prediction entropy value, the process of the processing device classifying the difficulty level of the sample query data in the sample query data set according to the prediction index value of each sample query data may include:
[0201] Similarly, the processing device can obtain the index value division ranges corresponding to the above-mentioned M difficulty levels, one difficulty level can correspond to one index value division range, and the M index value division ranges corresponding to the M difficulty levels are continuous with each other (i.e., the values are continuous with each other) and do not overlap with each other (i.e., different index value division ranges will not contain the same value). The index value division ranges corresponding to the M difficulty levels can be adaptively set in advance according to the actual business scenario. For example, the overall range of the predicted entropy value and the predicted loss value for predicting each sample query data according to the initial language model can be roughly evenly divided. The set M index value division ranges should be able to cover the predicted index value of each sample query data.
[0202] The processing device can calculate the sum of the predicted loss value and the predicted entropy value of the second sample query data, and can use the calculated sum (i.e., the sum of the predicted loss value and the predicted entropy value of the second sample query data) as the difficulty classification value of the second sample query data. The difficulty classification value can be a value used to divide the difficulty level of the second sample query data.
[0203] Alternatively, in another embodiment, corresponding weights may be set for the predicted loss value and the predicted entropy value, respectively. For example, a first weight may be set for the predicted loss value, and a second weight may be set for the predicted entropy value. Thus, the predicted loss value and the predicted entropy value of the second sample query data may be weighted and summed using the first and second weights to calculate the difficulty classification value of the second sample query data. For example, the product value of the first weight and the predicted loss value of the second sample query data may be calculated as the first product value, and the product value of the second weight and the predicted entropy value of the second sample query data may be calculated as the second product value. Thus, the sum of the first product value and the second product value may be used as the difficulty classification value of the second sample query data. The first weight and the second weight may be adaptively set according to actual business needs. For example, if the predicted loss value is considered more important than the predicted entropy value, the first weight may be set higher than the second weight. If the predicted entropy value is considered more important than the predicted loss value, the second weight may be set higher than the first weight.
[0204] The processing device can obtain the target index value division range within which the difficulty division value of the second sample query data falls within the aforementioned M index value division ranges, i.e., the difficulty division value of the second sample query data falls within the target index value division range. Therefore, the processing device can use the difficulty level corresponding to the target index value division range as the difficulty level of the second sample query data. For example, if the difficulty level corresponding to the target index value division range is the fifth difficulty level, then the second sample query data can be classified as sample query data of the fifth difficulty level.
[0205] For example, the predicted entropy and predicted loss values predicted by the initial language model for each sample query data, and the overall range of the difficulty classification values calculated for each sample query data are around 30 to 90. M is equal to 3, and the three difficulty levels can include, in sequence, a simple difficulty level, a medium difficulty level, and a high difficulty level. Then, the M index value classification ranges corresponding to the M difficulty levels can include, in sequence, [0, 50), [50, 70), and [70, +∞). The simple difficulty level can correspond to the index value classification range [0, 50), the medium difficulty level can correspond to the index value classification range [50, 70), and the high difficulty level can correspond to the index value classification range [70, +∞). If the difficulty classification value of the second sample query data is 35, then the target index value classification range for the difficulty classification value of the second sample query data is [0, 50). Therefore, the second sample query data can be classified into the simple difficulty level.
[0206] See Figure 6 , Figure 6 This is a schematic diagram of a scenario for classifying the difficulty level of the second sample query data provided by an embodiment of the present application. Figure 6 As shown, if the prediction index value of the second sample query data includes a predicted loss value and a predicted entropy value, the difficulty classification value of the second sample query data can be calculated by comprehensively calculating the predicted loss value and the predicted entropy value. Here, M can be equal to 3, and the three difficulty levels can include a simple difficulty level, a medium difficulty level, and a high difficulty level.
[0207] The easy difficulty level corresponds to the index value range f1, the medium difficulty level corresponds to the index value range f2, and the high difficulty level corresponds to the index value range f3. Therefore, if the difficulty level of the second sample query data falls within the index value range f2, the second sample query data can be classified into the medium difficulty level corresponding to the index value range f2.
[0208] The above is a specific explanation using the process of dividing the difficulty level of the second sample query data as an example. The processing device can divide the difficulty level of each sample query data in the sample query data set according to the same principle as the above-mentioned difficulty level division of the second sample query data, that is, divide each sample query data into its corresponding difficulty level.
[0209] Step S202 : selecting first sample query data based on sample query data of M difficulty levels.
[0210] In one embodiment, the present application can generally be based on sample query data of various difficulty levels, and according to the difficulty levels from easy to difficult (i.e., from low to high difficulty levels) into which the sample query data are divided, the initial language model can be trained step by step in a progressive manner. Therefore, when the initial language model is initially trained (i.e., when it is first trained), the sample query data of the first difficulty level (i.e., the simplest sample query data) among the sample query data of the M difficulty levels divided as described above can be used for training.
[0211] In the present application, after each round of iterative training of the initial language model, the effectiveness of the current round of training of the initial language model can be evaluated. Based on the effectiveness of the current round of training of the initial language model, sample query data for the next round of iterative training of the initial language model can be selected from the sample query dataset. That is, during the training of the initial language model, the difficulty level of the sample query data used in each round of iterative training of the initial language model can be dynamically changed based on the changes in the prediction accuracy of the initial language model for the sample query data during each round of iterative training. This prediction accuracy can be used to reflect the effectiveness of each round of training of the initial language model. For example, a higher prediction accuracy indicates a better training effect for the initial language model, and conversely, a lower prediction accuracy indicates a worse training effect for the initial language model.
[0212] Therefore, it can be understood that the first language model can be obtained after the Kth round of iterative training of the initial language model, where K is a non-negative integer, that is, the first language model can be a model after any round of iterative training of the initial language model. If K is equal to 0, it indicates that the initial language model has not been iteratively trained. Therefore, the first language model can be the initial language model. If K is a value greater than 0, such as 2, it indicates that the first language model can be obtained after the second round of iterative training of the initial language model. Therefore, the first language model can be an intermediate state in the training process of the initial language model (that is, the initial language model under training).
[0213] If K is 0, it indicates that the initial language model has not yet been trained with sample query data. Therefore, it can be assumed that the next round (i.e., the first round) will select sample query data of the first difficulty level from the sample query data of M difficulty levels to train the initial language model. The specific number of sample query data required for iterative training of the initial language model in each round (i.e., the number of sample query data used for one round of iterative training of the initial language model) can be set according to the actual application scenario, for example, 20 can be used.
[0214] If K is greater than 0, the process of selecting the first sample query data for the K+1th round of iterative training of the first language model using sample query data of M difficulty levels may include: the processing device may obtain the difficulty level of the sample query data used for model training in the Kth round of iterative training, and the difficulty level may be expressed as the i-th difficulty level, where i is a positive integer and i is less than or equal to M.
[0215] The processing device may also obtain the prediction accuracy of the initial language model during training for the sample query data during the Kth round of iterative training. This prediction accuracy may be referred to as the target prediction accuracy. The target prediction accuracy is the prediction accuracy of the initial language model during training for each sample query data during the Kth round of iterative training.
[0216] The following provides an exemplary method for obtaining the target prediction accuracy of sample query data during the Kth round of iterative training of the initial language model. In actual application scenarios, other methods can also be used to obtain the target prediction accuracy that can reflect the accuracy of the initial language model's prediction processing of sample query data during the Kth round of iterative training (i.e., the effect of the prediction processing).
[0217] Here, it is assumed that in the K-th round of iterative training, H sample query data are used to perform the K-th round of iterative training on the initial language model, and H is a positive integer. Therefore, the initial language model in training can predict N predicted response results for the H sample query data in the K-th round of iterative training, and one sample query data among the H sample query data can correspond to a group of N predicted response results. Moreover, the H sample query data can also each have their own corresponding reference response results, and one sample query data corresponds to one reference response result. In fact, similar to the first sample query data mentioned above, each sample query data in the sample query data set can have its own label information, and the label information of a sample query data is used to indicate the reference response result corresponding to the sample query data, and the reference response result is the ideal response result (i.e., the correct answer) for predicting the sample query data.
[0218] Therefore, the processing device can count the number of reference response results corresponding to each of the N predicted response results for H sample query data, and the H sample query data can correspond to a total of H counted quantities. Here, the counted quantity corresponding to a sample query data is the number of predicted response results that are consistent (i.e., identical) with the reference response result corresponding to the sample query data among the N predicted response results for the sample query data. This can be understood as the number of correct predictions made by the initial language model when performing prediction processing on the sample query data in the Kth round of iterative training.
[0219] The processing device can calculate the average number (i.e., the average value) between the H numbers, and can use the average number as the target prediction accuracy of the initial language model for the above H sample query data in the Kth round of iterative training. Alternatively, the processing device can also calculate the sum of the H numbers (i.e., summation), and can use the calculated sum value (i.e., summation number) as the target prediction accuracy. Among them, how to calculate the target prediction accuracy can be determined according to the actual application scenario, and this application does not limit this.
[0220] Alternatively, the processing device may also obtain the prediction loss values generated by the initial language model for the above-mentioned H sample query data in the Kth round of iterative training, and may sum the H prediction loss values corresponding to the H sample query data to obtain a summed loss value, and may use the inverse of the summed loss value or the value obtained by multiplying the inverse by an adjustment factor (which may be a constant set greater than 1) as the above-mentioned target prediction accuracy.
[0221] The above method of calculating the target prediction accuracy is only an exemplary description. In actual application scenarios, other feasible methods can also be adopted to calculate the target prediction accuracy.
[0222] The processing device can select the first sample query data for the K+1th round of iterative training of the first language model (i.e., the K+1th round of iterative training of the initial language model) from the sample query data of the M difficulty levels divided above based on the target prediction accuracy calculated above and the i-th difficulty level obtained (i.e., the difficulty level of the sample query data used to train the initial language model in the K-th round of iterative training), as described below.
[0223] If the above-mentioned target prediction accuracy meets the training effect improvement indication range corresponding to the initial language model, and i is not equal to M, the processing device can select the sample query data of the i+1th difficulty level from the sample query data of the M difficulty levels as the first sample query data. In other words, if the training effect of the initial language model is improved in the Kth round of iterative training (i.e., corresponding to the improvement of prediction accuracy), the difficulty of the sample query data used for the next round (i.e., the K+1th round) of iterative training of the initial language model can be increased. The training effect improvement indication range is the range set for evaluating the improvement of the training effect of the initial language model in one round of iterative training (which can belong to the range for prediction accuracy).
[0224] For example, the set training effect improvement indication range can be the range of [A, +∞], where +∞ represents positive infinity, A is a set constant, A can be a positive number, and the specific value of A can be set according to actual business needs. The processing device can calculate the prediction accuracy (which can be called a model effect evaluation parameter) obtained by subtracting the historical prediction accuracy from the target prediction accuracy. If the calculated prediction accuracy (i.e., the model effect evaluation parameter) is greater than or equal to A, it can be considered that the target prediction accuracy meets the training effect improvement indication range, indicating that the training effect of the initial language model has improved in the Kth round of iterative training. Therefore, the target prediction accuracy meeting the training effect improvement indication range can also be understood as: the model effect evaluation parameter calculated by the target prediction accuracy is within the training effect improvement indication range.
[0225] For example, the historical prediction accuracy can be the prediction accuracy of the sample query data when the initial language model undergoes the K-1th round of iterative training (i.e., the previous round of the current Kth round); or, if K is equal to 1, it indicates that there is no previous round of iterative training before the Kth round of iterative training. In this case, the historical prediction accuracy can be a preset prediction accuracy. The specific value of the preset prediction accuracy can be set according to actual business needs, and the preset prediction accuracy can be a slightly smaller value.
[0226] Alternatively, the historical prediction accuracy can also be the average prediction accuracy of the sample query data during the KG-th to K-1-th iterative training of the initial language model before the K-th iterative training. For example, the average prediction accuracy can be the average of G prediction accuracies of the sample query data during the KG-th to K-1-th iterative training of the initial language model, with one prediction accuracy corresponding to one iterative training. G is a positive integer, and the specific value of G can be set according to the actual application scenario, such as G can be set to 3. In other words, the average prediction accuracy of the G rounds of iterative training of the initial language model before the K-th iterative training can be used as the above-mentioned historical prediction accuracy. If there are no G rounds of iterative training before the K-th iterative training (i.e., less than G rounds of iterative training), the average prediction accuracy between each round of iterative training before the K-th iterative training can be used as the above-mentioned historical prediction accuracy. Similarly, if K is equal to 1, the above-mentioned preset prediction accuracy can also be used as the above-mentioned historical prediction accuracy.
[0227] If the target prediction accuracy meets the training effect improvement indication range corresponding to the initial language model, and i is equal to M, it indicates that the sample query data with the highest difficulty level has been used in the K-th round of iterative training. Therefore, the difficulty level of the sample query data used in the K+1-th iterative training can be kept unchanged, and the sample query data of the i-th difficulty level can continue to be selected from the sample query data of the M difficulty levels as the first sample query data.
[0228] Furthermore, if the target prediction accuracy meets the training effect stagnation indication range corresponding to the initial language model, the sample query data of the i-th difficulty level can be continuously selected from the sample query data of the M difficulty levels as the first sample query data. In other words, if the training effect of the initial language model has not been significantly improved in the K-th round of iterative training (i.e., the prediction accuracy has not been significantly improved), that is, the training effect is in a stagnant state, then the difficulty level of the sample query data for the next round of iterative training of the initial language model can be kept unchanged. The training effect stagnation indication range is the range set for evaluating that the training effect of the initial language model has not been significantly improved in one round of iterative training (it can also be a range for prediction accuracy).
[0229] For example, the set training effect stagnation indication range can be a range of [B, C], where both B and C can be set constants, B can be a negative number or a positive number, and C can be a positive number. The specific values of B and C can be set according to actual business needs. The processing device can calculate the prediction accuracy (i.e., the above-mentioned model effect evaluation parameter) obtained by subtracting the historical prediction accuracy from the target prediction accuracy. If the calculated prediction accuracy is greater than or equal to B and less than or equal to C, it can be considered that the target prediction accuracy meets the training effect stagnation indication range, indicating that the training effect of the initial language model has not been significantly improved in the Kth round of iterative training, and the model training effect has stagnated. Similarly, the target prediction accuracy meeting the training effect stagnation indication range can also be understood as: the model effect evaluation parameter calculated by the target prediction accuracy is within the training effect stagnation indication range.
[0230] And, if the target prediction accuracy meets the training effect degradation indication range corresponding to the initial language model, and i is not equal to 1, the sample query data of the i-1th difficulty level can be selected from the sample query data of the M difficulty levels as the first sample query data. In other words, if the training effect of the initial language model decreases in the Kth round of iterative training (i.e., corresponding to a decrease in prediction accuracy), the difficulty of the sample query data used for the next round (i.e., the K+1th round) of iterative training of the initial language model can be reduced to help the initial language model return to a suitable learning state. The training effect degradation indication range is the range set for evaluating the decline in the training effect of the initial language model in one round of iterative training (it can also be a range for prediction accuracy).
[0231] For example, the set training effect degradation indication range can be the range of [-∞, D], where -∞ represents negative infinity, and D can be a set constant. D can be a positive number or a negative number, and the specific value of D can be set according to actual business needs. The processing device can calculate the prediction accuracy (i.e., the above-mentioned model effect evaluation parameter) obtained by subtracting the above-mentioned historical prediction accuracy from the target prediction accuracy. If the calculated prediction accuracy is less than or equal to D, it can be considered that the target prediction accuracy meets the training effect degradation indication range, indicating that the training effect of the initial language model has declined in the Kth round of iterative training. Similarly, the target prediction accuracy meeting the training effect degradation indication range can also be understood as: the model effect evaluation parameter calculated by the target prediction accuracy is within the training effect degradation indication range.
[0232] If the target prediction accuracy falls within the training effect degradation indicator range corresponding to the initial language model, and i equals 1, then the sample query data of the i-th difficulty level can be selected from the sample query data of the M difficulty levels as the first sample query data. In other words, if the training effect of the initial language model degrades during the K-th round of iterative training, if the K-th round of iterative training was already training with the sample query data of the lowest difficulty level, then the sample query data of the first difficulty level can continue to be selected for model training during the K+1-th round of iterative training.
[0233] When selecting the first sample query data from the sample query data assigned to a certain difficulty level (e.g., the i-th difficulty level, the i+1-th difficulty level, or the i-1-th difficulty level), random selection may be employed, or other selection methods may be employed (e.g., preferentially selecting sample query data that has not been used for model training or has been used for model training the least number of times). The number of selected first sample query data may also be H, i.e., H sample query data may be used for model training in one round of iterative training.
[0234] In one embodiment, the training effect improvement indication range, training effect stagnation indication range, and training effect decline indication range can be continuous ranges (e.g., ranges with continuous values) and do not overlap (i.e., different ranges do not contain the same value), and the difficulty level of each sample query data used for model training in the same round of iterative training can be the same. The training effect improvement indication range, training effect stagnation indication range, and training effect decline indication range can all be set using corresponding thresholds, which can be reflected as the upper limit (i.e., maximum value) and lower limit (i.e., minimum value) of the training effect improvement indication range, training effect stagnation indication range, and training effect decline indication range.
[0235] Through the method provided in this application, the model can be trained stage by stage. For example, in the early stage of training, the initial language model can be trained with sample query data of a simple difficulty level, so that the initial language model can quickly master basic sample knowledge; in the middle stage of training, more challenging sample query data (i.e., sample query data with a higher difficulty level, such as sample query data with a medium difficulty level) can be introduced to continue training the initial language model, so that the initial language model can not only learn simple basic knowledge, but also learn more complex patterns and dependencies for predicting sample query data; and in the late stage of training, high-difficulty sample query data can be introduced to further train the initial language model, so that the initial language model can better generalize and solve the prediction problems of complex sample query data. This stage usually involves more noise and some reasoning patterns of sparser sample query data.
[0236] See Figure 7 , Figure 7 This is a schematic diagram of a scenario in which a first sample query data set is selected from a sample query data set provided by an embodiment of the present application. Figure 7 As shown, the sample query data set may include a large number of sample query data. Here, a sample query data set including Q sample query data is used as an example for explanation, where Q is a positive integer, that is, the sample query data set may include sample query data g1 to sample query data gQ.
[0237] For each sample query data in the sample query data set, the present application can generate a corresponding prediction index value for each sample query data through the initial language model, such as the prediction index value z1 generated for the sample query data g1 to the prediction index value zQ generated for the sample query data gQ. The present application can use the prediction index value generated for each sample query data to divide the sample query data in the sample query data set into difficulty levels, so as to divide the sample query data in the sample query data set into sample query data of M difficulty levels, including sample query data of the first difficulty level to sample query data of the Mth difficulty level.
[0238] Therefore, the present application can adaptively and dynamically select the first sample query data from the sample query data of M difficulty levels based on the performance of the current model training (such as the effect of the Kth round of iterative training) and the difficulty level of the currently used sample query data (such as the difficulty level of the sample query data used in the Kth round of iterative training) to achieve the next round of model iterative training.
[0239] Through the above process, it is achieved that when the initial language model is subjected to each round of iterative training, the sample query data used in each round of iterative training is dynamically selected according to the M difficulty levels. That is, when the initial language model is subjected to each round of iterative training, the difficulty level of the sample query data used can be dynamically and adaptively changed according to the training effect of the initial language model in each round of iterative training. That is, when the initial language model is subjected to each round of iterative training, the sample query data that meets the prediction ability of the initial language model currently being trained can be dynamically selected, thereby balancing the model training time and the selection of data difficulty, so as to ensure that the initial language model is trained step by step and improve the accuracy and reliability of the training of the initial language model. Since overly simple sample query data may cause the initial language model to overfit, and overly complex sample query data may cause the initial language model to fall into a local optimal solution, resulting in slow convergence, the present application dynamically selects sample query data of corresponding difficulty levels through the real-time training performance of the model to control the growth rate of the difficulty level of the sample query data (that is, the growth rate of the difficulty of the model training task), that is, to control the time point of introducing more complex model training tasks, thereby avoiding the occurrence of these problems.
[0240] This application combines the course learning optimization strategy and the reasoning process optimization based on answer feedback, which not only improves the training speed of the GRPO algorithm, but also makes the application of reinforcement learning in complex model training tasks (such as tasks that use sample query data with high difficulty levels for model training) more efficient, greatly improving the training effect of the model in complex model training tasks.
[0241] This application introduces an optimization strategy for course learning, and divides the training data (i.e., sample query data) into different batches according to difficulty for training, so that the model will first be trained on easy-to-solve training data (such as sample query data with a lower difficulty level), and gradually transition to more difficult training data. This approach not only reduces the training time required for the model on complex training data, but also avoids the model from getting into trouble too early when faced with difficult training data. Through this layered training, the model can gradually adapt to complex training data, improves the training efficiency of the model, and shows significant advantages in application scenarios of efficient reasoning and rapid convergence. This approach avoids the problems of slow learning speed (such as slow reasoning speed for more difficult training data), low learning efficiency (such as overfitting or underfitting problems), and poor learning effect caused by exposing the model to training data of various difficulties at the same time.
[0242] This application gradually increases the difficulty of the model training task so that the initial language model can learn complex data reasoning rules more efficiently. Especially when faced with large-scale and complex training data, the learning progress can be dynamically adjusted (that is, the difficulty of the sample query data for training) to effectively improve the learning ability of the initial language model during the training process and ensure the accuracy of the training of the initial language model.
[0243] See Figure 8 , Figure 8 This is a flow chart of a model training process provided by the embodiment of the present application. Figure 8 As shown, the process may include: 1. The present application may initialize tasks and model parameters. The initialization task may be a task for initializing model training, and the initialization model parameters may be used to obtain an initial language model. The initial language model has initialized model parameters, and the initial language model is the language model to be trained. 2. The present application may evaluate sample difficulty, that is, evaluate the difficulty (such as prediction difficulty) of sample query data in a sample query dataset, such as dividing the sample query data in the sample query dataset into difficulty levels according to the principles described above.
[0244] 3. The present application may select a model training task, which may be a task of selecting sample query data of a suitable difficulty level to continue iterative training of the initial language model currently being trained. For example, for the initial language model of the initial training, the selected model training task may be a task of selecting sample query data of the lowest difficulty level for model training, and for the initial language model in training, the selected model training task may be a task of selecting sample query data of a difficulty level that matches the training effect (i.e., training performance) of the initial language model currently being trained to perform model training. 4. The present application may perform iterative training of the model through the selected model training task. 5. The present application may evaluate the model performance after each round of iterative training of the initial language model, such as evaluating the model performance through the above-calculated prediction accuracy, and may adaptively select sample query data of the corresponding difficulty level for the next round of iterative model training based on the evaluated model performance, such as selecting to perform any one of the following steps 6 to 8 based on the evaluated model performance. Steps 6 to 8 may be three parallel steps.
[0245] 6. When the model performance improves (such as when the prediction accuracy of the model in the current round of iterative training meets the range indicating improved training effect), the difficulty of the model training task can be increased, that is, the difficulty level of the sample query data used for the next round of iterative training can be increased. 7. When the model performance stagnates (such as when the prediction accuracy of the model in the current round of iterative training meets the range indicating stagnation of training effect), the difficulty of the current model training task can be increased to maintain the same, or the difficulty of the model training task can be reduced (such as by one difficulty level), that is, the difficulty level of the sample query data used for the next round of iterative training can be kept unchanged or reduced. 8. When the model performance deteriorates (such as when the prediction accuracy of the model in the current round of iterative training meets the range indicating decreased training effect), the difficulty of the model training task can be reduced, that is, the difficulty level of the sample query data used for the next round of iterative training can be reduced.
[0246] 9. This application can repeat steps 3 to 8 above, and continuously perform multiple rounds of iterative training on the initial language model until the training stop condition is reached. At this time, the continued training of the initial language model can be terminated, and the trained second language model can be obtained.
[0247] By adopting the above-mentioned method of the present application, the initial language model can be strengthened by combining the selection of dynamic difficulty levels of sample query data during the training process and the direct provision of correct answers to the model when the model fails to predict the sample query data to guide the correct learning of the model. This not only greatly improves the efficiency of training the initial language model, but also ensures that the initial language model will not have the problem of failing to learn the correct prediction process, thereby ensuring the accuracy and reliability of the training of the initial language model.
[0248] The above-mentioned dynamic adjustment mechanism for the difficulty level of sample query data provided by this application can enable the model to maintain a high level of reasoning accuracy when facing unseen situations or uncertain problems. In addition, this application uses sample query data of various difficulty levels to gradually and progressively train the model, which can enhance the model's ability to handle diverse problems and improve the model's adaptability and generalization capabilities for various data reasoning tasks. In summary, the above-mentioned method provided by this application not only improves the accuracy and efficiency of the model in data prediction processing, but also enhances the model's adaptability, continuous learning ability, and real-time adjustment ability, thereby improving the overall performance of the model in data prediction processing.
[0249] See Figure 9 , Figure 9 This is a structural diagram of a data processing device provided in an embodiment of the present application. Figure 9 As shown, the data processing device 90 may include: an acquisition module 901 , a prediction module 902 , and a guidance module 903 .
[0250] An acquisition module 901 is configured to acquire a first language model and select first sample query data from a sample query dataset for training the first language model, wherein the first sample query data has label information indicating a target reference response result corresponding to the first sample query data.
[0251] Prediction module 902, configured to call a first language model to perform prediction processing on the first sample query data;
[0252] Guidance module 903 is used to guide the first language model to perform prediction processing on the first sample query data based on the target reference response result if the first language model fails to predict the first sample query data, thereby obtaining a second language model. The second language model is used to generate a corresponding prediction response result based on the input query data.
[0253] In one embodiment, the first language model generates N first predicted response results after performing prediction processing on the first sample query data, and each first predicted response result has its own prediction probability, where N is a positive integer;
[0254] The guidance module 903 guides the first language model to perform prediction processing on the first sample query data based on the target reference response result to obtain the second language model, including:
[0255] Determine the first predicted response result with the highest prediction probability among the N first predicted response results as the output response result;
[0256] Based on the output response result and the target reference response result, the first language model is guided to perform prediction processing on the first sample query data to obtain a second language model.
[0257] In one embodiment, the guidance module 903 guides the first language model to perform prediction processing on the first sample query data based on the output response result and the target reference response result to obtain the second language model, including:
[0258] Searching for negative knowledge information associated with the first sample query data from a preset knowledge graph based on the output response result; and
[0259] Searching the knowledge graph for positive knowledge information associated with the first sample query data based on the target reference response result;
[0260] Negative knowledge information and positive knowledge information are used to guide the first language model to perform prediction processing on the first sample query data to obtain a second language model.
[0261] In one embodiment, the first sample query data is prompt information of the first language model; the guidance module 903 uses the negative knowledge information and the positive knowledge information to guide the first language model to perform prediction processing on the first sample query data to obtain the second language model, including:
[0262] Using the negative knowledge information and the positive knowledge information as background information of the first sample query data to construct reconstruction prompt information for the first language model, where the reconstruction prompt information includes the first sample query data and the background information;
[0263] Inputting the reconstructed prompt information into the first language model, and calling the first language model to perform prediction processing on the reconstructed prompt information to generate an updated prediction response result for the first sample query data;
[0264] Based on the difference between the updated predicted response result and the target reference response result, the model parameters of the first language model are modified to obtain the second language model.
[0265] In one embodiment, the knowledge graph is constructed based on multiple entities and attribute information of multiple entities;
[0266] The guidance module 903 searches for negative knowledge information associated with the first sample query data from a preset knowledge graph based on the output response result, including:
[0267] Obtain the first entity associated with the output reply result, and search the knowledge graph for attribute information of the first entity as negative knowledge information;
[0268] Furthermore, the method in which the guiding module 903 searches the knowledge graph for positive knowledge information associated with the first sample query data based on the target reference reply result includes:
[0269] Obtain the second entity associated with the target reference reply result, and search the attribute information of the second entity from the knowledge graph as positive knowledge information.
[0270] In one embodiment, the guidance module 903 guides the first language model to perform prediction processing on the first sample query data based on the target reference response result to obtain the second language model, including:
[0271] inputting the target reference response result into the first language model;
[0272] The first language model is called to learn a prediction process of using the first sample query data to predict the target reference response result to obtain a second language model.
[0273] In one embodiment, the prediction process of the first language model learning to predict the target reference response result using the first sample query data is a process of inferring the correct prediction path for predicting the target reference response result using the first sample query data; and
[0274] The process of the first language model inferring the correct prediction path is a process of optimizing the model parameters of the first language model.
[0275] In one embodiment, the prediction module 902 uses the first language model to perform prediction processing on the first sample query data, including:
[0276] Calling the first language model to perform prediction processing on the first sample query data to generate N first prediction response results for the first sample query data;
[0277] If the N first predicted response results do not include the target reference response result, it is determined that the first language model fails to predict the first sample query data;
[0278] If the N first predicted response results include the target reference response result, it is determined that the first language model successfully predicts the first sample query data.
[0279] In one embodiment, the method in which the acquisition module 901 selects first sample query data for training the first language model from the sample query data set includes:
[0280] Performing difficulty classification on the sample query data in the sample query data set to obtain sample query data of M difficulty levels, where M is a positive integer;
[0281] Selecting first sample query data based on the sample query data of M difficulty levels;
[0282] The M difficulty levels include the first difficulty level to the Mth difficulty level, and the first difficulty level to the Mth difficulty level increase in sequence.
[0283] In one embodiment, the first language model is an initial language model or a language model obtained by training the initial language model;
[0284] The acquisition module 901 performs difficulty classification processing on the sample query data in the sample query data set to obtain the sample query data of M difficulty levels, including:
[0285] Calling the initial language model to perform prediction processing on each sample query data in the sample query data set to generate a prediction index value for each sample query data;
[0286] The sample query data in the sample query data set are divided into difficulty levels based on the prediction index value of each sample query data to obtain sample query data of M difficulty levels.
[0287] In one embodiment, any one of the sample query data sets is second sample query data, and the prediction indicator value of the second sample query data includes at least one of the following:
[0288] The prediction loss value of the initial language model for the second sample query data, and the prediction entropy value of the initial language model for the second sample query data;
[0289] The predicted entropy value refers to the entropy value of the predicted probability distribution generated by the initial language model for the second sample query data, and the predicted probability distribution is composed of the predicted probabilities of N predicted response results generated by the initial language model for the second sample query data.
[0290] In one embodiment, if the prediction index value of the second sample query data is a predicted loss value, the acquisition module 901 performs difficulty classification processing on the sample query data in the sample query data set based on the prediction index value of each sample query data to obtain sample query data of M difficulty levels, including:
[0291] Obtain loss value division ranges corresponding to M difficulty levels, where the M loss value division ranges corresponding to the M difficulty levels are continuous and non-overlapping;
[0292] Obtain a target loss value partition range in which the predicted loss value of the second sample query data lies among the M loss value partition ranges;
[0293] The difficulty level corresponding to the target loss value division range is determined as the difficulty level of the second sample query data.
[0294] In one embodiment, if the prediction index value of the second sample query data is a predicted entropy value, the acquisition module 901 performs difficulty classification processing on the sample query data in the sample query data set based on the prediction index value of each sample query data to obtain sample query data of M difficulty levels, including:
[0295] Obtaining entropy value division ranges corresponding to M difficulty levels, wherein the M entropy value division ranges corresponding to the M difficulty levels are continuous and non-overlapping;
[0296] Obtain a target entropy value partition range in which the predicted entropy value of the second sample query data lies in the M entropy value partition ranges;
[0297] The difficulty level corresponding to the target entropy value division range is determined as the difficulty level of the second sample query data.
[0298] In one embodiment, if the prediction index value of the second sample query data includes a prediction loss value and a prediction entropy value, the acquisition module 901 performs difficulty classification processing on the sample query data in the sample query data set based on the prediction index value of each sample query data to obtain sample query data of M difficulty levels, including:
[0299] Obtain index value division ranges corresponding to M difficulty levels, wherein the M index value division ranges corresponding to the M difficulty levels are continuous and non-overlapping;
[0300] Calculating the sum of the predicted loss value and the predicted entropy value of the second sample query data, and determining the calculated sum as the difficulty classification value of the second sample query data;
[0301] Obtain a target indicator value division range in which the difficulty division value of the second sample query data lies in the M indicator value division ranges;
[0302] The difficulty level corresponding to the target indicator value division range is determined as the difficulty level of the second sample query data.
[0303] In one embodiment, the first language model is an initial language model or a language model obtained by training the initial language model. When the initial language model is initially trained, sample query data of the first difficulty level among the sample query data of the M difficulty levels is used for training.
[0304] In the training process of the initial language model, the difficulty level of the sample query data used in each round of iterative training of the initial language model is dynamically changed according to the change in the prediction accuracy of the sample query data of the initial language model in each round of iterative training;
[0305] The first language model is obtained after the initial language model is trained for the Kth round of iterations, where K is a non-negative integer.
[0306] In one embodiment, the acquisition module 901 selects the first sample query data based on the sample query data of M difficulty levels, including:
[0307] Obtain the i-th difficulty level of the sample query data used in the K-th round of iterative training, where i is a positive integer and is less than or equal to M;
[0308] Obtain the target prediction accuracy of the initial language model in training for the sample query data during the Kth round of iterative training;
[0309] According to the target prediction accuracy and the i-th difficulty level, the first sample query data is selected from the sample query data of M difficulty levels.
[0310] In one embodiment, the initial language model in training predicts N predicted response results for each of the H sample query data in the Kth round of iterative training, and each of the H sample query data has its own reference response result, where H is a positive integer;
[0311] The acquisition module 901 acquires the target prediction accuracy of the initial language model in training for the sample query data during the Kth round of iterative training, including:
[0312] Count the number of reference response results in each of the N predicted response results for H sample query data, and the number of H corresponding reference response results for each of the H sample query data;
[0313] Calculate the average number among H numbers as the target prediction accuracy.
[0314] In one embodiment, the acquisition module 901 selects the first sample query data from the sample query data of M difficulty levels according to the target prediction accuracy and the i-th difficulty level, including:
[0315] If the target prediction accuracy meets the training effect improvement indication range corresponding to the initial language model, and i is not equal to M, then the sample query data of the i+1th difficulty level is selected from the sample query data of the M difficulty levels as the first sample query data;
[0316] If the target prediction accuracy meets the training effect improvement indication range, and i is equal to M, then the sample query data of the i-th difficulty level is selected from the sample query data of the M difficulty levels as the first sample query data;
[0317] If the target prediction accuracy meets the training effect stagnation indication range corresponding to the initial language model, then select the sample query data of the i-th difficulty level from the sample query data of the M difficulty levels as the first sample query data;
[0318] If the target prediction accuracy meets the training effect degradation indication range corresponding to the initial language model, and i is not equal to 1, then the sample query data of the i-1th difficulty level is selected from the sample query data of the M difficulty levels as the first sample query data;
[0319] If the target prediction accuracy meets the training effect degradation indication range, and i is equal to 1, sample query data of the i-th difficulty level is selected from the sample query data of the M difficulty levels as the first sample query data.
[0320] In one embodiment, the guidance module 903 is further configured to:
[0321] Obtaining query data sent by the client and inputting the query data into the second language model;
[0322] Calling the second language model to perform prediction processing on the input query data, generating N second predicted response results for the query data and a prediction probability of each second predicted response result;
[0323] Determine the second predicted response result with the highest prediction probability among the N second predicted response results as the target predicted response result for the query data;
[0324] The target prediction response result is returned to the client, so that the client outputs the target prediction response result.
[0325] According to one embodiment of the present application, Figure 3 The steps involved in the data processing method shown can be represented by Figure 9 The data processing apparatus 90 shown in FIG. Figure 3 The step S101 shown in FIG. Figure 9 The acquisition module 901 in is executed, Figure 3 The step S102 shown in FIG. Figure 9The prediction module 902 is executed; Figure 3 The step S103 shown in FIG. Figure 9 The guidance module 903 in is executed.
[0326] The present application can obtain a first language model and select first sample query data for training the first language model from a sample query data set, the first sample query data having label information, and the label information being used to indicate a target reference response result corresponding to the first sample query data; calling the first language model to perform prediction processing on the first sample query data; and, if the first language model fails to predict the first sample query data, guiding the first language model to perform prediction processing on the first sample query data based on the target reference response result, thereby obtaining a second language model, which can be used to generate a corresponding prediction response result based on the input query data. It can be seen that when the first language model fails to predict the first sample query data, the device proposed in the present application can directly guide the first language model to predict the first sample query data through the correct answer to the prediction processing of the first sample query data (that is, the target reference response result corresponding to the first sample query data), so that the first language model can quickly learn the process of correctly predicting the first sample query data from the correct direction, thereby obtaining a trained second language model, so that the first language model does not need to repeatedly predict the first sample query data multiple times, thereby improving the efficiency of training the second language model. Moreover, since the correct answer directly guides the first language model to learn the process of correctly predicting the first sample query data, the accuracy of the trained second language model is also guaranteed.
[0327] According to one embodiment of the present application, Figure 9 The various modules in the data processing device 90 shown can be individually or all combined into one or several units to constitute, or one (some) of the units can be further divided into multiple smaller sub-units in function, and the same operation can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In actual applications, the functions of a module can also be implemented by multiple units, or the functions of multiple modules can be implemented by one unit. In other embodiments of the present application, the data processing device 90 may also include other units. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0328] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0329] According to one embodiment of the present application, a computer program capable of executing the steps involved in the corresponding methods shown in the various embodiments of the present application can be run on a general-purpose computer device (the computer device may include processing elements and storage elements such as a central processing unit (CPU), a random access memory medium (RAM), and a read-only memory medium (ROM)) to construct the following. Figure 9 The data processing device 90 shown in FIG. The computer program may be recorded on a computer-readable recording medium, and may be loaded into the computer device through the computer-readable recording medium and executed therein.
[0330] See Figure 10 , Figure 10 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 10 As shown, the computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, in some embodiments, the computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 1005 may optionally be at least one storage device located away from the aforementioned processor 1001. As Figure 10 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a device control application program.
[0331] exist Figure 10In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an interface for user input; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:
[0332] Obtaining a first language model, and selecting first sample query data for training the first language model from a sample query dataset, wherein the first sample query data has label information, and the label information is used to indicate a target reference response result corresponding to the first sample query data;
[0333] Calling the first language model to perform prediction processing on the first sample query data;
[0334] If the first language model fails to predict the first sample query data, the first language model is guided to perform prediction processing on the first sample query data based on the target reference response result to obtain a second language model, which is used to generate a corresponding prediction response result based on the input query data.
[0335] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the above data processing method in each embodiment of the present application, and can also execute the above Figure 9 The description of the data processing device 90 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated here either.
[0336] In addition, it should be noted that the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program. When the processor executes the computer program, it can perform the description of the data processing method in each embodiment of the present application. Therefore, it will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated. For technical details not disclosed in the computer storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0337] The computer-readable storage medium may be an internal storage unit of the computer device, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may include both an internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0338] The present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device performs the description of the above-mentioned data processing method in each embodiment of the present application. Therefore, it will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0339] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other step units inherent to these processes, methods, apparatuses, products, or devices.
[0340] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0341] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Obtaining a first language model, and selecting first sample query data for training the first language model from a sample query dataset, wherein the first sample query data has label information, and the label information is used to indicate a target reference response result corresponding to the first sample query data; calling the first language model to perform prediction processing on the first sample query data; If the first language model fails to predict the first sample query data, the first language model is guided to perform prediction processing on the first sample query data based on the target reference response result to obtain a second language model, which is used to generate a corresponding prediction response result based on the input query data.
2. The method according to claim 1, wherein The first language model generates N first predicted response results after performing prediction processing on the first sample query data, and each first predicted response result has its own prediction probability, where N is a positive integer; The step of guiding the first language model to perform prediction processing on the first sample query data based on the target reference response result to obtain a second language model includes: Determine the first predicted response result with the highest prediction probability among the N first predicted response results as the output response result; Based on the output reply result and the target reference reply result, the first language model is guided to perform prediction processing on the first sample query data to obtain the second language model.
3. The method according to claim 2, wherein The step of guiding the first language model to perform prediction processing on the first sample query data based on the output response result and the target reference response result to obtain the second language model includes: Searching for negative knowledge information associated with the first sample query data from a preset knowledge graph based on the output reply result; and Searching the knowledge graph for positive knowledge information associated with the first sample query data based on the target reference response result; The negative knowledge information and the positive knowledge information are used to guide the first language model to perform prediction processing on the first sample query data to obtain the second language model.
4. The method according to claim 3, wherein The first sample query data belongs to prompt information of the first language model; and the using of the negative knowledge information and the positive knowledge information to guide the first language model to perform prediction processing on the first sample query data to obtain the second language model includes: Using the negative knowledge information and the positive knowledge information as background information of the first sample query data to construct reconstruction prompt information for the first language model, the reconstruction prompt information including the first sample query data and the background information; Inputting the reconstruction prompt information into the first language model, and calling the first language model to perform prediction processing on the reconstruction prompt information to generate an updated prediction response result for the first sample query data; Based on the difference between the updated predicted response result and the target reference response result, the model parameters of the first language model are modified to obtain the second language model.
5. The method according to claim 3, wherein The knowledge graph is constructed based on multiple entities and attribute information of the multiple entities; The step of searching a preset knowledge graph for negative knowledge information associated with the first sample query data based on the output reply result includes: Obtaining a first entity associated with the output reply result, and searching the knowledge graph for attribute information of the first entity as the negative knowledge information; Furthermore, searching the knowledge graph for positive knowledge information associated with the first sample query data based on the target reference reply result includes: Obtain a second entity associated with the target reference reply result, and search the knowledge graph for attribute information of the second entity as the positive knowledge information.
6. The method according to claim 1, wherein The step of guiding the first language model to perform prediction processing on the first sample query data based on the target reference response result to obtain a second language model includes: inputting the target reference response result into the first language model; The first language model is called to learn a prediction process of using the first sample query data to predict the target reference response result to obtain the second language model.
7. The method according to claim 1, wherein The selecting first sample query data for training the first language model from the sample query data set includes: Performing difficulty classification processing on the sample query data in the sample query data set to obtain sample query data of M difficulty levels, where M is a positive integer; selecting the first sample query data based on the sample query data of the M difficulty levels; The M difficulty levels include the first difficulty level to the Mth difficulty level in sequence, and the first difficulty level to the Mth difficulty level increase in sequence.
8. The method according to claim 7, wherein The first language model is an initial language model or a language model obtained by training the initial language model; The performing difficulty classification processing on the sample query data in the sample query data set to obtain sample query data of M difficulty levels includes: Calling the initial language model to perform prediction processing on each sample query data in the sample query data set to generate a prediction index value for each sample query data; The sample query data in the sample query data set are processed for difficulty classification based on the prediction index value of each sample query data to obtain the sample query data of the M difficulty levels.
9. The method according to claim 8, wherein Any one of the sample query data sets is second sample query data, and the prediction index value of the second sample query data includes at least one of the following: a prediction loss value of the initial language model for the second sample query data, and a prediction entropy value of the initial language model for the second sample query data; The predicted entropy value refers to the entropy value of the predicted probability distribution generated by the initial language model for the second sample query data, and the predicted probability distribution is composed of the predicted probabilities of N predicted response results generated by the initial language model for the second sample query data.
10. The method according to claim 9, wherein If the prediction index value of the second sample query data is the predicted loss value, performing difficulty classification processing on the sample query data in the sample query data set based on the prediction index value of each sample query data to obtain the sample query data of the M difficulty levels includes: Obtaining loss value division ranges corresponding to the M difficulty levels, respectively, where the M loss value division ranges corresponding to the M difficulty levels are continuous and non-overlapping; Obtaining a target loss value partition range in which the predicted loss value of the second sample query data lies in the M loss value partition ranges; The difficulty level corresponding to the target loss value division range is determined as the difficulty level of the second sample query data.
11. The method according to claim 9, wherein If the prediction index value of the second sample query data is the prediction entropy value, performing difficulty classification processing on the sample query data in the sample query data set based on the prediction index value of each sample query data to obtain the sample query data of the M difficulty levels includes: Obtaining entropy value division ranges corresponding to the M difficulty levels, respectively, where the M entropy value division ranges corresponding to the M difficulty levels are continuous and non-overlapping; Obtaining a target entropy value division range in which the predicted entropy value of the second sample query data lies in the M entropy value division ranges; The difficulty level corresponding to the target entropy value division range is determined as the difficulty level of the second sample query data.
12. The method according to claim 9, wherein If the prediction index value of the second sample query data includes the prediction loss value and the prediction entropy value, performing difficulty classification processing on the sample query data in the sample query data set based on the prediction index value of each sample query data to obtain the sample query data of the M difficulty levels includes: Obtaining index value division ranges corresponding to the M difficulty levels, respectively, where the M index value division ranges corresponding to the M difficulty levels are continuous and non-overlapping; Summing the predicted loss value and the predicted entropy value of the second sample query data, and determining the summed value as the difficulty classification value of the second sample query data; Obtaining a target index value division range in which the difficulty division value of the second sample query data lies in the M index value division ranges; The difficulty level corresponding to the target indicator value division range is determined as the difficulty level of the second sample query data.
13. The method according to claim 7, wherein The first language model is an initial language model or a language model obtained by training the initial language model, and when the initial language model is initially trained, the sample query data of the first difficulty level among the sample query data of the M difficulty levels is used for training; Wherein, during the training process of the initial language model, the difficulty level of the sample query data used in each round of iterative training of the initial language model is dynamically changed according to the change in the prediction accuracy of the sample query data of the initial language model during each round of iterative training; The first language model is obtained after performing the Kth round of iterative training on the initial language model, where K is a non-negative integer.
14. The method according to claim 13, wherein The selecting the first sample query data based on the sample query data of the M difficulty levels includes: Obtaining the i-th difficulty level of the sample query data used in the K-th round of iterative training, where i is a positive integer and is less than or equal to M; Obtaining a target prediction accuracy of the initial language model in training for sample query data during the K-th round of iterative training; The first sample query data is selected from the sample query data of the M difficulty levels according to the target prediction accuracy and the i-th difficulty level.
15. The method according to claim 14, wherein The initial language model in training predicts N predicted response results for each of the H sample query data in the K-th round of iterative training, and each of the H sample query data has its own reference response result, where H is a positive integer; The obtaining of a target prediction accuracy of the initial language model in training for sample query data during the K-th round of iterative training includes: Counting the number of reference response results corresponding to the N predicted response results of each of the H sample query data, where the H number of reference response results corresponds to the H sample query data; The average number among the H numbers is calculated as the target prediction accuracy.
16. The method according to claim 14, wherein The selecting the first sample query data from the sample query data of the M difficulty levels according to the target prediction accuracy and the i-th difficulty level includes: If the target prediction accuracy meets the training effect improvement indication range corresponding to the initial language model, and i is not equal to M, then selecting sample query data of the i+1th difficulty level from the sample query data of the M difficulty levels as the first sample query data; If the target prediction accuracy meets the training effect improvement indication range, and i is equal to M, then selecting sample query data of the i-th difficulty level from the sample query data of the M difficulty levels as the first sample query data; If the target prediction accuracy meets the training effect stagnation indication range corresponding to the initial language model, selecting sample query data of the i-th difficulty level from the sample query data of the M difficulty levels as the first sample query data; If the target prediction accuracy meets the training effect degradation indication range corresponding to the initial language model, and i is not equal to 1, selecting sample query data of the i-1th difficulty level from the sample query data of the M difficulty levels as the first sample query data; If the target prediction accuracy meets the training effect degradation indication range, and i is equal to 1, sample query data of the i-th difficulty level is selected from the sample query data of the M difficulty levels as the first sample query data.
17. A data processing device, characterized in that: The device comprises: an acquisition module, configured to acquire a first language model and select first sample query data from a sample query dataset for training the first language model, wherein the first sample query data has label information, and the label information is used to indicate a target reference response result corresponding to the first sample query data; a prediction module, configured to call the first language model to perform prediction processing on the first sample query data; A guidance module is configured to guide the first language model to perform prediction processing on the first sample query data based on the target reference response result if the first language model fails to predict the first sample query data, thereby obtaining a second language model, wherein the second language model is configured to generate a corresponding prediction response result based on the input query data.
18. A computer program product, characterized in that The computer program product comprises a computer program stored in a computer-readable storage medium. The computer program is suitable for being read and executed by a processor, so as to enable a computer device having the processor to perform the method according to any one of claims 1 to 16.
19. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 16.
20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the method according to any one of claims 1 to 16.