Model training method, natural language processing method and device
By generating forward and reverse reasoning data through the teacher's large language model, bidirectional training is performed on the student model, and the error data optimization strategy is used to solve the problem of insufficient model reasoning ability in natural language reasoning and improve the model's reasoning accuracy and self-reflection ability.
Patent Information
- Application Number
- CN202510878224.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies have limitations in the natural language reasoning process, making it difficult to effectively improve the model's reasoning ability and accuracy.
The forward and reverse reasoning data are generated by the teacher's large language model, and the student model is trained in both directions. The erroneous forward reasoning data and reverse reasoning data are combined, and the direct preference optimization strategy is used for model training to improve the model's reasoning ability.
It enhances the model's sensitivity to logical contradictions and semantic implicit relationships, reduces misjudgments, and improves the model's reasoning accuracy and self-reflection ability.
Smart Images

Figure CN120764685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a model training method, a natural language processing method and a device. Background Art
[0002] With the continuous development of technology, natural language processing has also made great progress. As natural language inference tasks become more complex, reasoning methods based on large language models have gradually become a research hotspot. The core of natural language inference lies in understanding the logical relationship between context and hypothesis and accurately classifying them accordingly.
[0003] However, related technologies still have certain limitations in the reasoning process. Therefore, how to perform natural language processing more effectively has become an urgent problem to be solved in the industry. Summary of the Invention
[0004] The present invention provides a model training method, a natural language processing method and a device to solve the problem of how to more effectively perform natural language processing in the prior art.
[0005] The present invention provides a model training method, comprising: Processing each natural language inference data through the teacher's large language model to obtain each forward inference data and each first reverse inference data; the natural language inference data includes: context information, hypothesis information, and category labels of the context information and the hypothesis information; Based on each of the forward reasoning data and each of the first reverse reasoning data, the student language model is trained to obtain a first language model with forward and reverse reasoning capabilities; Based on each target natural language inference data and the first large language model, constructing erroneous forward inference data corresponding to each target natural language inference data, and constructing second reverse inference data based on the erroneous forward inference data; based on the teacher large language model, processing the erroneous forward inference data corresponding to each target natural language inference data and the second reverse inference data to obtain target inference data with erroneous forward inference and correct reverse inference; According to each of the erroneous forward inference data and the target inference data, the first language model is trained by a direct preference optimization strategy to obtain a second language model.
[0006] According to a model training method provided by the present invention, each natural language inference data is processed by a teacher large language model to obtain each forward inference data and each first reverse inference data; comprising: Obtaining various natural language inference data; wherein the category labels in the natural language inference data include: support labels, contradictory labels, and neutral labels; Input each natural language reasoning data into the teacher's large language model to obtain a forward reasoning thinking description corresponding to each natural language reasoning data; and construct each forward reasoning data according to each natural language reasoning data and the forward reasoning thinking description corresponding to the natural language reasoning data; The forward reasoning data is input into the teacher's large language model, and the first reverse reasoning thinking description corresponding to each natural language reasoning data is output; based on each natural language reasoning data and the first reverse reasoning thinking description corresponding to the natural language reasoning data, each first reverse reasoning data is constructed.
[0007] According to a model training method provided by the present invention, the method trains a student's large language model based on each of the forward reasoning data and each of the first reverse reasoning data to obtain a first large language model with forward and reverse reasoning capabilities; the method comprises: Inputting the context information and hypothesis information in the forward reasoning data into the student language model, and outputting the forward category label prediction corresponding to the forward reasoning data and the forward reasoning thought description prediction; Inputting the context information and the category label in the first reverse reasoning data into the student language model, and outputting a hypothesis information prediction corresponding to the first reverse reasoning data and a first reverse reasoning thinking description prediction; Calculating a first cross entropy loss based on the forward category label prediction, the forward reasoning thought description prediction, and the category label and the forward reasoning thought description in the forward reasoning data; Calculating a second cross entropy loss based on the hypothesis information prediction, the first reverse reasoning thinking description prediction, and the hypothesis information and the first reverse reasoning thinking description in the first reverse reasoning data; The first cross entropy loss and the second cross entropy loss are used to optimize the student language model. The first reverse reasoning data and the forward reasoning data are traversed until a first preset training condition is met, thereby obtaining a first large language model with forward and reverse reasoning capabilities.
[0008] According to a model training method provided by the present invention, based on each target natural language inference data and the first large language model, constructing erroneous forward inference data corresponding to each target natural language inference data, and constructing second reverse inference data based on the erroneous forward inference data, including: Inputting each of the target natural language inference data into the first large language model to obtain a forward prediction error category label and an erroneous forward reasoning thought description for each of the target natural language inference data; Constructing each erroneous forward reasoning data based on each of the target natural language reasoning data, the forward prediction error category label corresponding to the target natural language reasoning data, and the erroneous forward reasoning thinking description; Based on the erroneous forward reasoning thinking description and the erroneous forward reasoning data, second reverse reasoning data is determined.
[0009] According to a model training method provided by the present invention, based on the erroneous forward reasoning thinking description and the erroneous forward reasoning data, determining the second reverse reasoning data includes: Based on the teacher's large language model, reverse reasoning is performed on the context information in the erroneous forward reasoning data and the forward predicted error category label to obtain a reverse prediction hypothesis from the context information to the forward predicted error category label, as well as a second reverse reasoning thinking description; Based on the context information, the forward prediction error category label, the reverse prediction hypothesis and the second reverse reasoning thinking description, second reverse reasoning data is constructed.
[0010] According to a model training method provided by the present invention, based on the teacher large language model, the erroneous forward reasoning data corresponding to each target natural language reasoning data and the second reverse reasoning data are processed to obtain target reasoning data with erroneous forward reasoning and correct reverse reasoning, including: Based on the teacher's large language model, the erroneous forward reasoning data and the second reverse reasoning data having the same context information are processed to obtain a predicted correct category label and a reflective reasoning thinking description of the predicted correct category label; According to the erroneous forward reasoning data and the second reverse reasoning data having the same context information, as well as the predicted correct category label and the reflective reasoning thinking description, target reasoning data in which the forward reasoning is erroneous and the reverse reasoning is correct is constructed.
[0011] The present invention also provides a natural language processing method, comprising: Obtain input context and hypothesis information; Inputting the context information and the hypothesis information into a second language model, outputting a category label between the context information and the hypothesis information, and a description of the reasoning thought process for obtaining the category label; The second largest language model is obtained based on any of the above-mentioned model training methods.
[0012] The present invention also provides a natural language processing method, wherein the reasoning thinking description includes: after a forward reasoning error, performing reverse reasoning to obtain the content of the category label.
[0013] The present invention also provides a model training device, comprising: A first processing module is configured to process each natural language inference data using a teacher large language model to obtain each forward inference data and each first reverse inference data; the natural language inference data includes: context information, hypothesis information, and category labels of the context information and the hypothesis information; A first training module is used to train the student's large language model based on each of the forward reasoning data and each of the first reverse reasoning data to obtain a first large language model with forward and reverse reasoning capabilities; A construction module, configured to construct, based on each target natural language inference data and the first language model, erroneous forward inference data corresponding to each target natural language inference data, and second reverse inference data; A second processing module is used to process the erroneous forward reasoning data and the second reverse reasoning data corresponding to each of the target natural language reasoning data based on the teacher large language model to obtain target reasoning data with erroneous forward reasoning and correct reverse reasoning; The second training module is used to train the first language model according to each of the erroneous forward inference data and the target inference data through a direct preference optimization strategy to obtain a second language model.
[0014] The present invention also provides a natural language processing device, comprising: The acquisition module is used to obtain input context information and hypothesis information; a third processing module, configured to input the context information and the hypothesis information into a second language model, and output a category label between the context information and the hypothesis information, and a description of the reasoning process for obtaining the category label; The second largest language model is obtained based on any of the above-mentioned model training methods.
[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any one of the above-mentioned model training methods or the natural language processing method.
[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned model training methods or natural language processing methods.
[0017] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any one of the above-mentioned model training methods or natural language processing methods.
[0018] The model training method, natural language processing method and device provided by the present invention first process the natural language reasoning data through the teacher model to obtain forward reasoning data and first reverse reasoning data. The forward reasoning data can help the model to forward deduce the hypothesis based on the context information, while the first reverse reasoning data reversely deduce the hypothesis-related information based on the category label and context information. The combination of the two provides training materials in both positive and negative directions for the student model. The student model is then trained based on these two-way data sets, so that the student model pays attention to reverse verification while learning forward deduction, understands and captures the logical relationship between context and hypothesis from different angles, enhances sensitivity to logical contradictions and semantic implicit relationships, reduces misjudgments, and obtains the first language model with forward and reverse reasoning capabilities. Next, based on each target natural language inference data and the first large language model, incorrect forward inference data and second reverse inference data corresponding to each target natural language inference data are constructed. The incorrect forward inference data and second reverse inference data are processed using the teacher large language model to obtain target inference data with incorrect forward inference and correct reverse inference. Based on the incorrect forward inference data and the target inference data, the first large language model is trained using a direct preference optimization strategy to obtain the second large language model. The direct preference optimization strategy guides the model to learn the correct inference preferences by comparing incorrect inference paths with correct inference paths, enabling the model to dynamically reflect on and correct its own inference paths, effectively improving the model's reasoning and natural language processing capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 It is a flow chart of the model training method provided by the present invention; Figure 2 A flow chart of the natural language processing method provided by the present invention; Figure 3 This is a schematic diagram of the structure of the model training device provided by the present invention; Figure 4 A schematic diagram of the structure of the natural language processing device provided by the present invention; Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0022] Figure 1 It is a flow chart of the model training method provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 110: Processing each natural language inference data using the teacher's large language model to obtain each forward inference data and each first backward inference data; the natural language inference data includes: context information, hypothesis information, and category labels of the context information and the hypothesis information; In this paper, natural language inference data is used to train and evaluate natural language inference models. Each natural language inference data sample contains contextual information, hypothesis information, and the category labels between them. These data samples provide the model with the necessary inputs and corresponding output labels for inference, enabling the model to learn how to classify hypotheses based on contextual information.
[0023] Contextual information is one of the fundamental input components in natural language inference tasks. It can be a paragraph of text, a sentence, or a combination of sentences, providing the background or factual statements associated with the hypothesis. Contextual information provides the starting point and basis for the model's reasoning. The model needs to extract key information from it to understand its relationship to the hypothesis. For example, in natural language inference tasks involving news reports and commentary, the content of the news report serves as contextual information, from which the model needs to extract factual details to make judgments about the hypothesis in the commentary.
[0024] Hypothesis information is another important input component in natural language reasoning tasks. It is typically a statement or proposition that needs to be verified, associated with contextual information. It can be an inference, assumption, or opinion about the context. The model's task is to determine the correctness or rationality of this hypothesis based on the context. For example, in product review analysis, the hypothesis might be "This product is very useful." The model needs to consider the context (such as the user's detailed description of the user experience) to determine whether this hypothesis is true.
[0025] In this invention, category labels are annotations of the relationship between context information and hypothesis information in natural language inference data, which indicate the type of logical relationship between the two. Specifically, they are divided into: Entailment: When the contextual information is sufficient to infer that the hypothesis is true, the category label is "support." For example, if the context is "I went to the library yesterday and borrowed several books on artificial intelligence" and the hypothesis is "I am interested in artificial intelligence," common sense suggests that borrowing books on related topics is likely due to interest in the topic, so the category label is "support."
[0026] Contradiction: If the context clearly indicates that the hypothesis is wrong, the class label is "contradiction." For example, if the context is "It's sunny today" and the hypothesis is "It's raining today," sunny and rainy are opposite weather conditions. The context contradicts the hypothesis, so the class label is "contradiction."
[0027] Neutral: A class label is neutral when the context neither supports nor disproves the hypothesis. For example, the context is "I bought a book" and the hypothesis is "I like to read novels." While buying a book may be related to reading, it doesn't guarantee that the book I bought is a novel, nor does it guarantee that I like to read novels. Therefore, the relationship between the two is neutral.
[0028] In this paper, the teacher language model is a pre-trained large language model with strong language understanding and generation capabilities for natural language processing tasks. In this process, the teacher model acts as a "mentor," processing natural language inference data to generate forward and backward inference data, providing guidance and reference for the training of the student model. The excellent performance of the teacher model ensures that the generated data is of high quality for training the student model.
[0029] In the present invention, forward reasoning data Contains contextual information , hypothesis information , category labels and forward reasoning thinking description ; First reverse reasoning data This includes contextual information , category labels , hypothesis information And the first description of reverse reasoning thinking .
[0030] In the present invention, positive thinking describes It is generated by the teacher's large language model based on the contextual information, hypothesis information, and category labels in the natural language inference data. Specifically, when the category label is known, forward reasoning is performed based on the contextual information and hypothesis information to confirm the reasoning content of the category label.
[0031] In the present invention, the first reverse thinking description is generated by the teacher's large language model based on the context information, hypothesis information, and category labels in the natural language reasoning data. Specifically, when the hypothesis information is known, the reasoning content of the hypothesis corresponding to the category label is reversely deduced based on the context information and category label.
[0032] Optionally, the first reverse thinking description can also be obtained by the teacher's large language model based on the context information, hypothesis information, and forward thinking description in the forward thinking description. Specifically, it can be that when the hypothesis information is known, based on the context information and category label, further reference is made to the forward thinking description to reversely deduce the reasoning content of the hypothesis corresponding to the category label.
[0033] Step 120: training the student language model based on each of the forward reasoning data and each of the first reverse reasoning data to obtain a first language model with forward and reverse reasoning capabilities; In the present invention, the student language model has been pre-trained before the training begins and has a certain foundation for language understanding and generation, but its forward and reverse reasoning capabilities need to be further strengthened.
[0034] In the present invention, the forward reasoning data and the first backward reasoning data are the input basis of step 120. Forward reasoning data Contains contextual information , hypothesis information , category labels and forward reasoning thinking description ; First reverse reasoning data This includes contextual information , category labels , hypothesis information And the first description of reverse reasoning thinking .
[0035] The forward inference data and the first backward inference data are integrated to form a training data set. The data is preprocessed by cleaning, normalizing, and other operations to ensure data quality.
[0036] Input context information and hypothesis information into the student's large language model to capture key features in the context and hypothesis and their interrelationships. Based on the forward reasoning data, the model predicts category labels and generates forward reasoning thought description predictions, and outputs the forward category label predictions corresponding to the forward reasoning data and the forward reasoning thought description predictions. Based on the first reverse reasoning data, the category label and context information are input into the student language model, the reverse hypothesis is predicted and a reverse reasoning thinking description is generated, and the hypothesis information prediction corresponding to the first reverse reasoning data and the first reverse reasoning thinking description prediction are output; Further, calculating a first cross entropy loss based on the forward category label prediction, the forward reasoning thought description prediction, and the category label and the forward reasoning thought description in the forward reasoning data; Calculating a second cross entropy loss based on the hypothesis information prediction, the first reverse reasoning thinking description prediction, and the hypothesis information and the first reverse reasoning thinking description in the first reverse reasoning data; Based on the first and second cross-entropy losses, a joint loss function is determined, and the model parameters are updated using the AdamW optimizer. The AdamW optimizer incorporates a weight decay mechanism based on an adaptive learning rate to prevent overfitting and improve model generalization. In each iteration, the gradient is calculated and the model parameters are updated, gradually reducing the value of the joint loss function.
[0037] After multiple rounds of iterative training, when the joint loss function converges or reaches the preset maximum number of iterations, the parameters of the student model are optimized and the forward and reverse reasoning capabilities are significantly improved, thus obtaining the first language model with forward and reverse reasoning capabilities.
[0038] Step 130: constructing error forward inference data corresponding to each target natural language inference data based on each target natural language inference data and the first language model, and constructing second reverse inference data based on the error forward inference data; In the present invention, the first large language model is the above-mentioned student large language model, which is trained based on forward reasoning data and first reverse reasoning data. The trained first large language model has forward reasoning ability and reverse reasoning ability.
[0039] In the present invention, the target natural language inference data may refer to data having the same data type as the above-mentioned natural language inference data but different data content.
[0040] Specifically, the target natural language inference data may also include: context information, hypothesis information, and category labels of the context information and the hypothesis information; however, the data content of the target natural language inference data is different from that of the above-mentioned natural language inference data, so as to bring more diverse data to the model and improve the model training effect.
[0041] In the present invention, after the target natural language reasoning data is input into the first largest language model, the first largest language model will output the forward prediction category label corresponding to each target natural language reasoning data, as well as the forward reasoning thinking description prediction; further, the category label carried in the target natural language reasoning data can be compared with the forward prediction category label output by the first largest language model. If the category label is consistent with the forward prediction category label, it is judged that the output forward prediction category label is a positive correct prediction category label, and the corresponding forward reasoning thinking description prediction is a correct forward reasoning thinking description prediction.
[0042] However, if the category label carried in the target natural language inference data is inconsistent with the positive prediction category label output by the first language model, it is judged that the output positive prediction category label is a positive error prediction category label. , the corresponding forward reasoning thinking description prediction is the wrong forward reasoning thinking description prediction .
[0043] In the present invention, the above method can be used to traverse each target natural language inference data to determine the positive error prediction category label output by the first language model. , and the corresponding incorrect forward reasoning thinking description prediction Based on the contextual information in the target natural language , hypothesis information , and the corresponding positive error prediction class label , incorrect forward reasoning thinking describes prediction , construct error forward inference data ( , , , ).
[0044] Each error forward inference data ( , , , ) includes: context information , hypothesis information , positive error prediction class label , The description of the wrong forward reasoning thinking .
[0045] In the present invention, the positive misprediction class label is the misprediction label made by the first language model for the hypothesis. For example, if the first language model mistakenly predicts a hypothesis as "support" when it should be "contradictory", then the misprediction class label is "support".
[0046] The description of incorrect forward reasoning describes the model's reasoning process when it arrives at an incorrect prediction label, reflecting the model's faulty reasoning logic at the time. For example, the model might mistakenly believe that an irrelevant detail mentioned in the context supports the hypothesis, and this faulty reasoning process will be reflected in the description.
[0047] In the present invention, the second reverse reasoning data It is generated by the teacher's large language model based on erroneous forward inference data and is used to help the model correct erroneous data.
[0048] The second reverse reasoning data specifically includes: context information , positive error prediction class label , reverse prediction hypothesis , and the second reverse reasoning description ; reverse prediction hypothesis It refers to the new hypothesis generated by the teacher's large language model based on contextual information and positive error prediction category labels. This hypothesis is designed to guide the model to learn from errors and understand what a reasonable hypothesis should be under the context and error category labels.
[0049] The backward reasoning description describes the reasoning process when the teacher's large language model generates a backward prediction hypothesis, explaining why this backward prediction hypothesis is generated. This description provides the model with the correct reasoning ideas and logic, helping the model understand how to reflect on its mistakes and draw correct conclusions.
[0050] For example, the context information of the target natural language inference data is "I borrowed a book on artificial intelligence from the library", the hypothesis information is "I am interested in artificial intelligence", and the category label is "support".
[0051] After inputting the contextual information and hypothesis information into the first language model, the positive prediction category label output by the model is "neutral", and the corresponding positive reasoning thinking description prediction is "the user just borrowed the book and did not explicitly express interest in artificial intelligence."
[0052] Based on the comparison between the forward prediction category label "neutral" and the category label "support" in the target natural language inference data, it can be determined that the forward prediction category label is a positive error prediction category label, and correspondingly, its forward reasoning thinking description prediction is also an incorrect forward reasoning thinking description prediction.
[0053] The constructed incorrect forward inference data includes contextual information "I borrowed a book on artificial intelligence from the library"; hypothesis information "I am interested in artificial intelligence"; incorrect positive prediction category label "neutral" and incorrect forward inference thinking description "the user only borrowed the book and did not explicitly express interest in artificial intelligence."
[0054] Then, the positive error prediction category label ("neutral") and contextual information ("I borrowed the AI book from the library") are input into the teacher's large language model for reverse reasoning.
[0055] After analysis, the teacher model concluded that judging interest solely based on borrowing behavior was inadequate. Therefore, it generated a reverse prediction hypothesis: "The user may have some interest in AI." It also output a reverse reasoning statement: "Borrowing AI books suggests the user may have some interest in AI. Although not explicitly expressed, such behavior is generally associated with interest. Therefore, the reverse reasoning supports the category label 'support'." These results will be used in subsequent model optimization training to help the model learn how to make more reasonable inferences in similar situations.
[0056] Step 140: Based on the teacher's large language model, the erroneous forward reasoning data and the second reverse reasoning data corresponding to each target natural language reasoning data are processed to obtain target reasoning data with erroneous forward reasoning and correct reverse reasoning. In the present invention, the teacher's large language model For pairs of incorrect forward inference data with the same context information and the second reverse reasoning data , and the context information Corresponding category labels Processing is performed to obtain the corresponding predicted category label, and how to perform reflective reasoning to obtain a reflective reasoning description of the predicted category label.
[0057] Predicting class labels with contextual information Corresponding category labels If the predicted category label is consistent, the predicted category label is determined to be the correct category label , whose reflective reasoning is described as .
[0058] Target inference data ( ) retains the original context information , hypothesis information and category labels , also includes the positive error prediction class label , incorrect forward reasoning thinking description reverse prediction hypothesis , and the second reverse reasoning description , and the predicted correct class labels generated by the teacher model , reflective reasoning description .
[0059] In the present invention, the target reasoning data shows the comparison between forward reasoning errors and backward reasoning correctness, as well as the reflection content, which helps the model learn how to reflect on the errors of forward reasoning and find the correct reasoning path.
[0060] In this invention, by comparing incorrect forward reasoning with correct reasoning, the correct reasoning logic and thinking mode are learned. By learning these target reasoning data, the model can improve the accuracy and rationality of its reasoning and reduce similar errors.
[0061] In an optional embodiment, the teacher large language model comprehensively analyzes the erroneous forward reasoning data and the second backward reasoning data.
[0062] The incorrect forward inference data includes contextual information "I borrowed a book on artificial intelligence from the library", hypothesis information "I am interested in artificial intelligence", incorrect positive prediction category label "neutral" and incorrect forward inference thinking description "the user only borrowed the book and did not explicitly express interest in artificial intelligence".
[0063] The second reverse reasoning data includes contextual information, the wrong positive prediction category label "neutral", the reverse prediction hypothesis "the user may have a certain interest in artificial intelligence" and the reverse reasoning thinking description "borrowing artificial intelligence books suggests that the user may have a certain interest in artificial intelligence. Although it is not explicitly expressed, such behavior is usually related to interest, so the reverse reasoning agrees with the category label as 'support'."
[0064] The incorrect forward reasoning data contains the model's incorrect output from forward reasoning, while the second reverse reasoning data provides the correct reasoning path from the reverse reasoning perspective based on the same context and incorrect category label. By comparing these two types of data, the teacher's large language model can clearly identify incorrect patterns in forward reasoning and correct logic in reverse reasoning, providing a basis for correcting errors.
[0065] Therefore, the incorrect forward inference data and the second backward inference data were fed into the teacher's large language model. The teacher's large language model compared the forward predicted category label "neutral" in the incorrect forward inference data with the backward inference result of the supporting category label "support" in the second backward inference data, confirming that the forward inference was incorrect. The teacher's large language model generated a reflective reasoning description: "Borrowing AI-related books usually indicates that the user has a certain interest in AI. Although the user does not explicitly express this, such behavior is related to this interest. Therefore, the correct category label should be 'support.' However, the previous forward inference did not fully consider the implicit relationship between borrowing behavior and interest, resulting in an incorrect judgment." The teacher large language model determines that the predicted correct class label for the target natural language inference data should be "support".
[0066] Step 150 : Based on each of the erroneous forward inference data and the target inference data, the first language model is trained by a direct preference optimization strategy to obtain a second language model.
[0067] In the present invention, the second largest language model is obtained by training the first largest language model obtained by the above training, using the erroneous forward reasoning data and the target reasoning data through a direct preference optimization strategy.
[0068] Specifically, in the present invention, the student's large language model is first trained based on the forward reasoning data and the first reverse reasoning data to obtain the first large language model, and then the first large language model is further trained for the second time based on the erroneous forward reasoning data and the target reasoning data through a direct preference optimization strategy to obtain the second large language model.
[0069] In the present invention, the direct preference optimization strategy is a reinforcement learning strategy that enables the model to learn to prefer the correct reasoning path by comparing positive and negative sample pairs.
[0070] Specifically, the wrong forward inference data is used as a negative sample, and the target inference data is used as a positive sample. A positive and negative sample pair dataset is constructed based on the wrong forward inference data and the target inference data. , the model conducts comparative learning on the data set through positive and negative samples, and adjusts the parameters to increase the probability of generating similar outputs to positive samples, while reducing the probability of generating similar outputs to negative samples.
[0071] In this direct preference optimization, both the training model and the baseline model serve as the primary language model. The training model is the model to be optimized, while the baseline model is used for comparison, helping the training model learn the correct reasoning preferences. By comparing the performance of the training model and the baseline model on positive and negative sample pairs, the model can learn the correct reasoning path, thereby improving the accuracy and rationality of reasoning.
[0072] More specifically, the AdamW optimizer is used to optimize the large language model Perform direct preference optimization and calculate the joint loss function To update the model parameters until the maximum number of iterations or the loss function is reached Until convergence, a large language model with self-reflection and forward and reverse reasoning capabilities is obtained, which is used to predict the relationship between the input context and the hypothesis, and obtain the corresponding thought process description and its predicted label.
[0073] In the specific implementation, the loss function As shown in the following formula: ; in, is the Sigmoid function, used to smooth the learning results; is a hyperparameter that controls the strength of preference learning; Represents the model to be trained, Represents the baseline model. In this direct preference optimization, the training model and the baseline model both use the first language model. , to enhance The performance of the model, Indicates the treatment of training model Given input Get the output The corresponding probability of Indicates the treatment of training model Given input Get the output The corresponding probability of Represents the baseline model Given input Get the output The corresponding probability of Represents the baseline model Given input Get the output The corresponding probability of .
[0074] In the present invention, after training, a second-largest language model is obtained, which has stronger self-reflection ability and forward and reverse reasoning ability. It can identify erroneous forward reasoning data, reflect on it, and finally correct it to obtain the correct prediction results based on the reflection data. It can more accurately predict the relationship between context and hypothesis, and output the corresponding thinking process description and prediction label.
[0075] In the present application, the natural language reasoning data is first processed by the teacher model to obtain forward reasoning data and first reverse reasoning data. The forward reasoning data can help the model to derive the hypothesis from the context information, while the first reverse reasoning data derives the hypothesis related information from the category label and the context information. The combination of the two provides the student model with training materials in both forward and reverse directions. Then, based on these bidirectional data sets, the student model is trained to learn forward derivation while focusing on reverse verification, understand and capture the logical relationship between context and hypothesis from different angles, enhance the sensitivity to logical contradictions and semantic implicit relationships, reduce misjudgment, and obtain the first large language model with forward and reverse reasoning capabilities. Subsequently, based on the natural language reasoning data and the first large language model, error forward reasoning data corresponding to each target natural language reasoning data and second reverse reasoning data are constructed; the teacher large language model is used to process the error forward reasoning data and the second reverse reasoning data to obtain target reasoning data with error forward reasoning and correct reverse reasoning. According to the error forward reasoning data and the target reasoning data, the first large language model is trained using a direct preference optimization strategy to obtain a second large language model. The direct preference optimization strategy compares the error reasoning path and the correct reasoning path to guide the model to learn the correct reasoning preference, so that the model can dynamically reflect and correct its reasoning path, effectively improving the reasoning ability and natural language processing ability of the model.
[0076] Optionally, the processing of each natural language reasoning data by the teacher large language model to obtain each forward reasoning data and each first reverse reasoning data comprises: Each natural language reasoning data is obtained; wherein the category label in the natural language reasoning data comprises a support label, a contradiction label and a neutral label; Each natural language reasoning data is input into the teacher large language model to obtain the corresponding forward reasoning thought description of each natural language reasoning data. According to each natural language reasoning data and the corresponding forward reasoning thought description of the natural language reasoning data, each forward reasoning data is constructed. The forward reasoning data is input into the teacher large language model to output the first reverse reasoning thought description corresponding to each natural language reasoning data. According to each natural language reasoning data and the corresponding first reverse reasoning thought description of the natural language reasoning data, each first reverse reasoning data is constructed.
[0077] In the present application, the forward reasoning thought description is generated by the teacher large language model according to the context information and hypothesis information in the natural language reasoning data, representing the first A forward reasoning thought description, which reflects the detailed thinking process of reasoning forward from the context to the hypothesis.
[0078] The forward reasoning thought description usually includes the predicted category label of the hypothesis by the model and the detailed description of the reasoning process. These descriptions are used to train the forward reasoning ability of the student model, enabling the student model to learn how to accurately predict the category label of the hypothesis based on the context information.
[0079] Based on natural language reasoning data and corresponding forward reasoning thought descriptions , forward reasoning data are constructed. The forward reasoning data not only contains the original context information , hypothesis information and category label , but also adds the forward reasoning thought description generated by the teacher model . These data are used to train the forward reasoning ability of the student model, enabling the student model to learn how to accurately predict the category label of the hypothesis based on the context information and understand the reasoning process.
[0080] In the present application, the first reverse reasoning thought description is generated by the teacher large language model based on the forward reasoning data, represents the ith first reverse reasoning thought description, which reflects the thinking process of reasoning backward from the category label and the context information. It describes how the model reverses the hypothesis consistent with the category label from the known category label and context information, including the logic of reverse reasoning, evidence extraction and reasoning steps, etc.
[0081] In the present application, the teacher large language model not only provides guidance for forward reasoning for the student model, but also generates examples of reverse reasoning to help the student model understand and learn natural language reasoning tasks from different directions. These data will be used to train the student model to have forward and reverse reasoning ability.
[0082] Optionally, the student large language model is trained based on each of the forward reasoning data and each of the first reverse reasoning data to obtain a first large language model with forward and reverse reasoning ability; comprising: inputting the context information and the hypothesis information in the forward reasoning data into the student large language model, outputting the predicted forward category label and the predicted forward reasoning thought description corresponding to the forward reasoning data; inputting the context information and the category label in the first reverse reasoning data into the student large language model, outputting the predicted hypothesis information and the predicted first reverse reasoning thought description corresponding to the first reverse reasoning data; Calculating a first cross entropy loss based on the forward category label prediction, the forward reasoning thought description prediction, and the category label and the forward reasoning thought description in the forward reasoning data; Calculating a second cross entropy loss based on the hypothesis information prediction, the first reverse reasoning thinking description prediction, and the hypothesis information and the first reverse reasoning thinking description in the first reverse reasoning data; The first cross entropy loss and the second cross entropy loss are used to optimize the student language model. The first reverse reasoning data and the forward reasoning data are traversed until a first preset training condition is met, thereby obtaining a first large language model with forward and reverse reasoning capabilities.
[0083] In the present invention, the context information in the forward reasoning data is and hypothesis information After inputting the student language model, the model will output the corresponding positive category label prediction and forward reasoning to describe predictions .
[0084] In the present invention, the context information in the first reverse reasoning data is and the category labels Input the student language model and output the hypothesis information prediction corresponding to the first reverse reasoning data , and the first reverse reasoning thinking describes the prediction ; According to the positive class label prediction , forward reasoning thinking describes prediction , and the category labels in the forward inference data , forward reasoning thinking description , calculate the first cross entropy loss : ; Predictions based on the hypothesis information , the first reverse reasoning thinking describes the prediction , and the hypothesis information in the first reverse reasoning data , the first reverse reasoning thinking description , calculate the second cross entropy loss ; ; Constructing a joint loss function : in, Represents the weight control parameters of forward and reverse learning, which is used to force the model to learn forward and reverse thinking at the same time.
[0085] In the present invention, the first preset training condition may specifically refer to reaching a predetermined number of iterations, the performance of the model on the validation set no longer significantly improving, the training loss falling below a certain threshold, etc.
[0086] When the first preset training condition is met, it is considered that the student's large language model has been fully trained and has a certain forward and reverse reasoning ability.
[0087] The trained model parameters and related model structures are saved to form the first language model with forward and reverse reasoning capabilities. The first language model can perform forward and reverse reasoning at the same time in new natural language reasoning tasks based on given context and hypothesis information, providing more comprehensive and accurate judgment results.
[0088] Optionally, based on each target natural language inference data and the first language model, constructing erroneous forward inference data corresponding to each target natural language inference data and second reverse inference data includes: Inputting each of the target natural language inference data into the first large language model to obtain a forward prediction error category label and an erroneous forward reasoning thought description for each of the target natural language inference data; Constructing each erroneous forward reasoning data based on each of the target natural language reasoning data, the forward prediction error category label corresponding to the target natural language reasoning data, and the erroneous forward reasoning thinking description; Based on the erroneous forward reasoning thinking description and the erroneous forward reasoning data, second reverse reasoning data is determined.
[0089] In this paper, the process of inputting target natural language inference data into the first language model to obtain forward prediction error category labels and descriptions of incorrect forward inference thinking is a key feedback link in model training. It is used to identify and record model errors during the forward inference process, providing key data support for subsequent model optimization.
[0090] Specifically, each target natural language inference data contains contextual information and hypothesis information. Each target natural language inference data is input into the first language model. The model performs forward reasoning based on the input contextual information and hypothesis information, and outputs the corresponding forward prediction category label and forward reasoning thinking description.
[0091] The category label predicted by the model is compared with the true category label. If the prediction is wrong, the wrong category label is recorded to obtain the positive prediction error category label. , the wrong forward reasoning thinking description refers to when the forward prediction category label is identified as wrong, the corresponding forward reasoning thinking description is the wrong forward reasoning thinking description , which reflects the model's reasoning process when making incorrect predictions.
[0092] In the present invention, each error forward inference data ( , , , ) includes: context information , hypothesis information , positive error prediction class label , Description of the incorrect forward reasoning thinking .
[0093] Optionally, determining second reverse reasoning data based on the incorrect forward reasoning thinking description and the incorrect forward reasoning data includes: Based on the teacher's large language model, reverse reasoning is performed on the context information in the erroneous forward reasoning data and the forward predicted error category label to obtain a reverse prediction hypothesis from the context information to the forward predicted error category label, as well as a second reverse reasoning thinking description; Based on the context information, the forward prediction error category label, the reverse prediction hypothesis and the second reverse reasoning thinking description, second reverse reasoning data is constructed.
[0094] In the present invention, the erroneous forward reasoning data and the second reverse reasoning data are input into the teacher language model. These two types of data share the same context information but contain different reasoning directions and reasoning results.
[0095] The teacher's large language model uses its powerful language understanding and reasoning capabilities to analyze the incorrect forward reasoning thinking descriptions in the incorrect forward reasoning data and identify the problems in the reasoning process. At the same time, it refers to the correct reverse reasoning path in the second reverse reasoning data, that is, the reverse prediction hypothesis and reverse reasoning thinking description in the second reverse reasoning data.
[0096] Combined with the analysis of incorrect forward reasoning and the reference of correct reverse reasoning, the teacher model predicts the correct category label Make predictions and generate corresponding reflective reasoning descriptions .
[0097] Reflective reasoning description The derivation process from incorrect forward inference to the correct category label is explained, reflecting the teacher model's correction of errors and guidance of the correct reasoning path.
[0098] Target inference data Preserves the original context information , hypothesis information and category labels , also includes the positive error prediction class label , reverse prediction hypothesis , and the second reverse reasoning description , and the predicted correct class labels generated by the teacher model , reflective reasoning description .
[0099] In the present invention, the target reasoning data shows the comparison between forward reasoning errors and reverse reasoning correctness, which helps the model learn how to reflect on the errors in forward reasoning and find the correct reverse reasoning path.
[0100] Figure 2 The flow chart of the natural language processing method provided by the present invention is as follows: Figure 2 Shown, including: Step 210, obtaining input context information and hypothesis information; Step 220: Input the context information and the hypothesis information into a second language model, and output a category label between the context information and the hypothesis information, as well as a description of the reasoning process used to obtain the category label. Among them, the second largest language model is obtained based on the above model training method.
[0101] In the present invention, in natural language reasoning tasks, context information and hypothesis information are the basis for the model to perform reasoning. Context information provides the background and prerequisites for reasoning, while hypothesis information is the statement or proposition that needs to be verified.
[0102] The obtained contextual information and hypothesis information are input into the trained second largest language model.
[0103] The second largest language model has been fully trained through the previous steps and has strong forward and reverse reasoning and self-reflection capabilities.
[0104] The second largest language model outputs the category label between the context information and the hypothesis information, as well as a description of the reasoning thinking that inferred the category label.
[0105] The category label represents the logical relationship between the two, such as "support," "contradiction," or "neutral." The reasoning description details the model's reasoning process, including how it starts from the context and the logical steps it takes to ultimately reach the conclusion for the category label.
[0106] For example, the second language model can be used to verify news facts. Contextual information can refer to the specific content of a news report, including details such as the time, location, people involved, and the course of events. Hypothesis information refers to various statements or opinions circulating in society about the news.
[0107] The news report and the rumor are fed into a second language model, which then outputs a category label. If the model outputs a "support" label, it indicates the rumor aligns with the news report and is likely true and reliable. If it outputs a "contradictory" label, it indicates the rumor contradicts the news report and is likely a rumor or false information. A "neutral" label indicates that the truth or falsity of the rumor cannot be determined based on the available information.
[0108] The second language model in this invention can help the public quickly distinguish the authenticity of information, curb the spread of rumors, and increase the public's trust in the news media.
[0109] In another embodiment, the second language model can also be applied to a customer service system. Context information refers to the conversation history between the user and the customer service representative, including the user's problem description, previous customer service representative's response, and related product or service information. Hypothesis information refers to the user's new problem or need. When a user asks a new question, the conversation history and the new question are fed into the second largest language model. Based on the previous conversation content and the user's new question, the model outputs the relationship between the two. If the model determines that the new question is related to the previous conversation content and can be answered based on the existing information, it outputs the corresponding answer. If it determines that further inquiry or transfer to human customer service is required, it provides corresponding suggestions.
[0110] The solution of the present invention can improve customer service efficiency and user satisfaction, reduce the workload of manual customer service, and achieve more efficient and intelligent customer service.
[0111] The present invention first processes natural language inference data through a teacher model to obtain forward inference data and first backward inference data. The forward inference data helps the model to forward-derive hypotheses based on contextual information, while the first backward inference data reversely infers hypothesis-related information based on category labels and contextual information. The combination of the two provides training material for the student model in both forward and backward directions. The student model is then trained based on these bidirectional data sets, allowing the student model to learn forward inference while focusing on backward verification. This allows the student model to understand and capture the logical relationship between context and hypothesis from different perspectives, enhances sensitivity to logical contradictions and semantic implicit relationships, reduces misjudgments, and obtains a first language model with forward and backward inference capabilities. Subsequently, based on the natural language inference data and the first language model, incorrect forward inference data and second backward inference data corresponding to each target natural language inference data are constructed. The incorrect forward inference data and the second backward inference data are processed using the teacher language model to obtain target inference data with incorrect forward inference and correct backward inference. Based on the incorrect forward inference data and the target inference data, a direct preference optimization strategy is used to train the first language model to obtain a second language model. The direct preference optimization strategy guides the model to learn the correct reasoning preferences by comparing incorrect reasoning paths with correct reasoning paths, enabling the model to dynamically reflect on and correct its own reasoning paths, effectively improving the model's reasoning ability and natural language processing capabilities.
[0112] Optionally, the reasoning thinking description includes: after a forward reasoning error, reverse reasoning and reflection are performed to obtain the content of the category label.
[0113] In the present invention, when the second language model makes an error during forward reasoning, it identifies inconsistencies or contradictions between the reasoning results and the contextual information and hypothesis. At this time, the model will initiate reflection and backward reasoning mechanisms.
[0114] During the reflection and backward reasoning process, the second language model attempts to construct logical relationships consistent with possible correct category labels, identifying key details or semantic relationships that may have been overlooked in the previous forward reasoning and adjusting its understanding of the hypothesis accordingly. At the same time, the model corrects previous incorrect reasoning paths and re-evaluates the logical connection between the hypothesis and the context to ensure that the new reasoning process supports the correct category label.
[0115] Ultimately, through reflection and backward reasoning, the second-largest language model generates a new reasoning path that reasonably explains why the correct category label applies to the current context and hypothesis. The reasoning description obtained in this process details how the model reflects on the errors in forward reasoning and finds the correct answer through backward reasoning.
[0116] The model training device and natural language processing device provided by the present invention are described below. The device described below and the model training method and natural language processing method described above can be referenced to each other.
[0117] Figure 3 This is a schematic diagram of the structure of the model training device provided by the present invention, as shown in FIG. Figure 3 Shown, including: The first processing module 310 is configured to process each natural language inference data using the teacher's large language model to obtain each forward inference data and each first backward inference data; the natural language inference data includes: context information, hypothesis information, and category labels of the context information and the hypothesis information; The first training module 320 is used to train the student's large language model based on each of the forward reasoning data and each of the first reverse reasoning data to obtain a first large language model with forward and reverse reasoning capabilities; The construction module 330 is used to construct the error forward reasoning data corresponding to each target natural language reasoning data and the second reverse reasoning data based on each target natural language reasoning data and the first language model; The second processing module 340 is used to process the erroneous forward reasoning data and the second reverse reasoning data corresponding to each of the target natural language reasoning data based on the teacher large language model to obtain target reasoning data with erroneous forward reasoning and correct reverse reasoning; The second training module 350 is used to train the first language model according to each of the erroneous forward inference data and the target inference data through a direct preference optimization strategy to obtain a second language model.
[0118] Figure 4 A schematic diagram of the structure of the natural language processing device provided by the present invention is shown in FIG. Figure 4 Shown, including: The acquisition module 410 is used to obtain input context information and hypothesis information; The third processing module 420 is configured to input the context information and the hypothesis information into the second largest language model, and output a category label between the context information and the hypothesis information, and a description of the reasoning process for obtaining the category label; Among them, the second largest language model is obtained based on the above model training method.
[0119] In the present invention, the natural language inference data is first processed by the teacher model to obtain forward inference data and first reverse inference data. The forward inference data can help the model to forward deduce the hypothesis based on contextual information, while the first reverse inference data reversely deduce hypothesis-related information based on category labels and contextual information. The combination of the two provides training materials for the student model in both positive and negative directions. The student model is then trained based on these bidirectional data sets, so that the student model focuses on reverse verification while learning forward inference, understanding and capturing the logical relationship between context and hypothesis from different perspectives, enhancing sensitivity to logical contradictions and semantic implicit relationships, reducing misjudgments, and obtaining a first large language model with forward and reverse inference capabilities. Afterwards, based on the natural language inference data and the first large language model, incorrect forward inference data and second reverse inference data corresponding to each target natural language inference data are constructed; the incorrect forward inference data and second reverse inference data are processed using the teacher large language model to obtain target inference data with incorrect forward inference and correct reverse inference. Based on the incorrect forward inference data and the target inference data, the first large language model is trained using a direct preference optimization strategy to obtain a second large language model. The direct preference optimization strategy guides the model to learn the correct reasoning preferences by comparing incorrect reasoning paths with correct reasoning paths, enabling the model to dynamically reflect on and correct its own reasoning paths, effectively improving the model's reasoning ability and natural language processing capabilities.
[0120] Figure 5 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute a model training method or a natural language processing method, which includes: processing various natural language inference data using a teacher language model to obtain various forward inference data and various first reverse inference data; the natural language inference data includes: context information, hypothesis information, and category labels of the context information and the hypothesis information; Based on each of the forward reasoning data and each of the first reverse reasoning data, the student language model is trained to obtain a first language model with forward and reverse reasoning capabilities; Based on each target natural language inference data and the first language model, constructing error forward inference data corresponding to each target natural language inference data and second reverse inference data; Based on the teacher's large language model, the erroneous forward reasoning data and the second reverse reasoning data corresponding to each of the target natural language reasoning data are processed to obtain target reasoning data with erroneous forward reasoning and correct reverse reasoning; According to each of the erroneous forward inference data and the target inference data, the first language model is trained by a direct preference optimization strategy to obtain a second language model; Or, obtain input context information and hypothesis information; The context information and the hypothesis information are input into a second language model, and a category label between the context information and the hypothesis information and a description of the reasoning thinking used to obtain the category label are output.
[0121] Furthermore, the logic instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0122] On the other hand, the present invention further provides a computer program product, comprising a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the model training method or natural language processing method provided by the above methods, the method comprising: processing each natural language inference data using a teacher language model to obtain each forward inference data and each first backward inference data; the natural language inference data comprising: context information, hypothesis information, and category labels for the context information and the hypothesis information; Based on each of the forward reasoning data and each of the first reverse reasoning data, the student language model is trained to obtain a first language model with forward and reverse reasoning capabilities; Based on each target natural language inference data and the first language model, constructing error forward inference data corresponding to each target natural language inference data and second reverse inference data; Based on the teacher's large language model, the erroneous forward reasoning data and the second reverse reasoning data corresponding to each of the target natural language reasoning data are processed to obtain target reasoning data with erroneous forward reasoning and correct reverse reasoning; According to each of the erroneous forward inference data and the target inference data, the first language model is trained by a direct preference optimization strategy to obtain a second language model; Or, obtain input context information and hypothesis information; The context information and the hypothesis information are input into a second language model, and a category label between the context information and the hypothesis information and a description of the reasoning thinking used to obtain the category label are output.
[0123] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the model training method or natural language processing method provided by the above methods, the method comprising: processing each natural language inference data using a teacher language model to obtain each forward inference data and each first backward inference data; the natural language inference data comprising: context information, hypothesis information, and category labels for the context information and the hypothesis information; Based on each of the forward reasoning data and each of the first reverse reasoning data, the student language model is trained to obtain a first language model with forward and reverse reasoning capabilities; Based on each target natural language inference data and the first language model, constructing error forward inference data corresponding to each target natural language inference data and second reverse inference data; Based on the teacher's large language model, the erroneous forward reasoning data and the second reverse reasoning data corresponding to each of the target natural language reasoning data are processed to obtain target reasoning data with erroneous forward reasoning and correct reverse reasoning; According to each of the erroneous forward inference data and the target inference data, the first language model is trained by a direct preference optimization strategy to obtain a second language model; Or, obtain input context information and hypothesis information; The context information and the hypothesis information are input into a second language model, and a category label between the context information and the hypothesis information and a description of the reasoning thinking used to obtain the category label are output.
[0124] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0125] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A model training method, characterized in that: include: Through the teacher's large language model, each natural language inference data is processed to obtain each forward inference data and each first reverse inference data; The natural language inference data includes: context information, hypothesis information, and category labels of the context information and the hypothesis information; Based on each of the forward reasoning data and each of the first reverse reasoning data, the student language model is trained to obtain a first language model with forward and reverse reasoning capabilities; Based on each target natural language inference data and the first large language model, constructing erroneous forward inference data corresponding to each target natural language inference data, and constructing second reverse inference data based on the erroneous forward inference data; based on the teacher large language model, processing the erroneous forward inference data corresponding to each target natural language inference data and the second reverse inference data to obtain target inference data with erroneous forward inference and correct reverse inference; According to each of the erroneous forward inference data and the target inference data, the first language model is trained by a direct preference optimization strategy to obtain a second language model.
2. The model training method according to claim 1, characterized in that The method processes each natural language inference data through the teacher's large language model to obtain each forward inference data and each first reverse inference data; including: Obtaining various natural language inference data; wherein the category labels in the natural language inference data include: support labels, contradictory labels, and neutral labels; Input each natural language reasoning data into the teacher's large language model to obtain a forward reasoning thinking description corresponding to each natural language reasoning data; and construct each forward reasoning data according to each natural language reasoning data and the forward reasoning thinking description corresponding to the natural language reasoning data; The forward reasoning data is input into the teacher's large language model, and the first reverse reasoning thinking description corresponding to each natural language reasoning data is output; based on each natural language reasoning data and the first reverse reasoning thinking description corresponding to the natural language reasoning data, each first reverse reasoning data is constructed.
3. The model training method according to claim 1, characterized in that The method of training the student's large language model based on each of the forward reasoning data and each of the first reverse reasoning data to obtain a first large language model with forward and reverse reasoning capabilities includes: Inputting the context information and hypothesis information in the forward reasoning data into the student language model, and outputting the forward category label prediction corresponding to the forward reasoning data and the forward reasoning thought description prediction; Inputting the context information and the category label in the first reverse reasoning data into the student language model, and outputting a hypothesis information prediction corresponding to the first reverse reasoning data and a first reverse reasoning thinking description prediction; Calculating a first cross entropy loss based on the forward category label prediction, the forward reasoning thought description prediction, and the category label and the forward reasoning thought description in the forward reasoning data; Calculating a second cross entropy loss based on the hypothesis information prediction, the first reverse reasoning thinking description prediction, and the hypothesis information and the first reverse reasoning thinking description in the first reverse reasoning data; The first cross entropy loss and the second cross entropy loss are used to optimize the student language model. The first reverse reasoning data and the forward reasoning data are traversed until a first preset training condition is met, thereby obtaining a first large language model with forward and reverse reasoning capabilities.
4. The model training method according to claim 1, characterized in that Based on each target natural language inference data and the first language model, constructing error forward inference data corresponding to each target natural language inference data, and constructing second reverse inference data based on the error forward inference data; including: Inputting each of the target natural language inference data into the first large language model to obtain a forward prediction error category label and an erroneous forward reasoning thought description for each of the target natural language inference data; Constructing each erroneous forward reasoning data based on each of the target natural language reasoning data, the forward prediction error category label corresponding to the target natural language reasoning data, and the erroneous forward reasoning thinking description; Based on the erroneous forward reasoning thinking description and the erroneous forward reasoning data, second reverse reasoning data is determined.
5. The model training method according to claim 4, characterized in that Determining second reverse reasoning data based on the incorrect forward reasoning thinking description and the incorrect forward reasoning data includes: Based on the teacher's large language model, reverse reasoning is performed on the context information in the erroneous forward reasoning data and the forward predicted error category label to obtain a reverse prediction hypothesis from the context information to the forward predicted error category label, as well as a second reverse reasoning thinking description; Based on the context information, the forward prediction error category label, the reverse prediction hypothesis and the second reverse reasoning thinking description, second reverse reasoning data is constructed.
6. The model training method according to claim 5, characterized in that Based on the teacher's large language model, the erroneous forward reasoning data and the second reverse reasoning data corresponding to each of the target natural language reasoning data are processed to obtain target reasoning data with erroneous forward reasoning and correct reverse reasoning, including: Based on the teacher's large language model, the erroneous forward reasoning data and the second reverse reasoning data having the same context information are processed to obtain a predicted correct category label and a reflective reasoning thinking description of the predicted correct category label; According to the erroneous forward reasoning data and the second reverse reasoning data having the same context information, as well as the predicted correct category label and the reflective reasoning thinking description, target reasoning data in which the forward reasoning is erroneous and the reverse reasoning is correct is constructed.
7. A natural language processing method, characterized in that: include: Obtain input context and hypothesis information; Inputting the context information and the hypothesis information into a second language model, outputting a category label between the context information and the hypothesis information, and a description of the reasoning thought process for obtaining the category label; The second largest language model is obtained based on the model training method described in any one of claims 1 to 6.
8. The natural language processing method according to claim 7, characterized in that: The reasoning thinking description includes: after the forward reasoning is wrong, reverse reasoning and reflection are performed to obtain the content of the category label.
9. A model training device, characterized in that: include: A first processing module is used to process each natural language inference data through the teacher's large language model to obtain each forward inference data and each first reverse inference data; The natural language inference data includes: context information, hypothesis information, and category labels of the context information and the hypothesis information; A first training module is used to train the student's large language model based on each of the forward reasoning data and each of the first reverse reasoning data to obtain a first large language model with forward and reverse reasoning capabilities; A construction module is used to construct, based on each target natural language inference data and the first large language model, erroneous forward inference data corresponding to each target natural language inference data, and to construct second reverse inference data based on the erroneous forward inference data; a second processing module is used to process, based on the teacher large language model, the erroneous forward inference data corresponding to each target natural language inference data and the second reverse inference data to obtain target inference data in which the forward inference is erroneous and the reverse inference is correct; The second training module is used to train the first language model according to each of the erroneous forward inference data and the target inference data through a direct preference optimization strategy to obtain a second language model.
10. A natural language processing device, characterized in that: include: The acquisition module is used to obtain input context information and hypothesis information; a third processing module, configured to input the context information and the hypothesis information into a second language model, and output a category label between the context information and the hypothesis information, and a description of the reasoning process for obtaining the category label; The second largest language model is obtained based on the model training method described in any one of claims 1 to 6.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the model training method according to any one of claims 1 to 6, or the natural language processing method according to any one of claims 7-8.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the model training method according to any one of claims 1 to 6, or the natural language processing method according to any one of claims 7-8.
Citation Information
Cited By
Natural language programming method and device, electronic equipment and storage medium
CN122086375A
Natural language programming method and device, electronic equipment and storage medium
CN122086375B