Large language model common sense reasoning enhancement method and device based on knowledge filtering
Through a knowledge filtering-based method, the common sense reasoning ability of large language models is enhanced, and the problem of existing models performing poorly in common sense reasoning tasks is solved, achieving higher quality thinking chain generation and more accurate common sense reasoning answers.
Patent Information
- Application Number
- CN202510135729.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
The performance of existing large language models in common sense reasoning tasks is still far from that of human abilities, and it is difficult to effectively improve the quality of their generated thinking chains and the accuracy of answering common sense reasoning tasks.
Through a knowledge filtering method, a data set including common sense knowledge questions and answers, construct a question-knowledge pair, evaluate the degree of influence of knowledge on answers, generate level tags, adjust the parameters of the preset reward model, and realize the reasoning enhancement of the large language model.
The quality of the generated thinking chain of large language models and the accuracy of answering common sense reasoning tasks are significantly improved, while maintaining a good balance between positive effects and damage to the original reasoning ability.
Smart Images

Figure CN120069073A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of natural language processing technology, and more specifically, to a method and apparatus for enhancing commonsense reasoning of a large language model based on knowledge filtering. Background Art
[0002] In related fields, commonsense reasoning is one of the key capabilities in the field of artificial intelligence (abbreviated as AI), aiming to imitate the ability of humans to reason based on common sense and experience in daily life. This ability is crucial for AI systems because it can help AI systems better understand and handle problems in complex situations, thereby improving their decision-making ability and task execution effect. To evaluate this ability, commonsense reasoning tasks can be designed, and these commonsense reasoning tasks require the model to answer questions based on commonsense knowledge (for example, Figure 3 the shown examples). In recent research, compared with small models, large language models (abbreviated as large language models or large models (LLMs)) have shown improved performance in this task. However, the gap with human actual ability is still large.
[0003] Therefore, to solve the problems existing in related fields, an improved model is needed to address the above issues. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method and apparatus for enhancing reasoning of a large language model based on knowledge filtering. By using knowledge filtering, the commonsense reasoning ability of the large language model is enhanced, greatly improving the quality of the thought chain generated by the large language model and the accuracy of answering commonsense reasoning tasks.
[0005] In one general aspect, a method for enhancing reasoning of a large language model based on knowledge filtering is provided. The reasoning enhancement method includes: obtaining a data set including a plurality of questions and corresponding answers for commonsense knowledge; inputting the plurality of questions into a large language model based on knowledge filtering to obtain corresponding commonsense knowledge, so as to construct a first question-knowledge pair; classifying the first question-knowledge pair and generating a corresponding level label by evaluating the influence degree of the corresponding commonsense knowledge on the corresponding answers of the plurality of questions, to obtain a second question-knowledge pair data including the level label; inputting the second question-knowledge pair data and the data set into a preset reward model to obtain a corresponding reward value output; and adjusting each parameter of the preset reward model according to the reward value output to obtain a trained reward model, so as to achieve the enhancement of the reasoning of the large language model based on knowledge filtering.
[0006] Optionally, the inference enhancement method may further include: obtaining an inference problem related to common sense knowledge; inputting the inference problem into the large language model based on knowledge filtering to obtain corresponding multiple inference chains; inputting the corresponding multiple inference chains into the trained reward model to obtain a valid inference chain corresponding to the inference problem, which is used as the inference chain after inference enhancement.
[0007] Optionally, the inference enhancement method may further include: inputting the inference problem and the inference chain after inference enhancement into the large language model based on knowledge filtering to obtain multiple answers to the inference problem after multiple inferences, and using the answer with the highest marginal probability among the multiple answers as the final answer to the inference problem, thereby obtaining the inference answer after inference enhancement.
[0008] Optionally, the step of classifying the first question-knowledge pair and generating a corresponding level label based on the influence degree of the corresponding common sense knowledge on the corresponding answers to multiple questions, so as to obtain the second question-knowledge pair data including the level label may include: inputting the data set and the first question-knowledge pair into the large language model based on knowledge filtering to obtain a first answer and a second answer, where the first answer is obtained when the large language model based on knowledge filtering uses the corresponding common sense knowledge to solve the answer, and the second answer is obtained when the large language model based on knowledge filtering does not use the corresponding common sense knowledge to solve the answer; evaluating the influence degree of the corresponding common sense knowledge on the corresponding answers to multiple questions according to the first answer and the second answer, classifying the first question-knowledge pair and generating a corresponding level label, so as to obtain the second question-knowledge pair data including the level label.
[0009] Optionally, the step of adjusting each parameter of the preset reward model according to the output of the reward value to obtain the trained reward model may include: calculating a loss function value for adjusting each parameter through the following formula according to the output of the reward value:
[0010] L(θ) = -ylog(f(q, k; θ)) - (1 - y)log(1 - f(q, k; θ)),
[0011] where θ represents the model parameters of the preset reward model, L represents the loss function, y represents the level label, f represents the reward value predicted by the preset reward model, q represents the question, and k represents the corresponding common sense knowledge; adjusting each parameter based on the loss function value to obtain the trained reward model.
[0012] In another general aspect, there is provided an inference enhancement device for a large language model based on knowledge filtering. The inference enhancement device includes: a data acquisition module configured to acquire a data set including a plurality of questions regarding common sense knowledge and corresponding a plurality of answers; a data construction module configured to input the plurality of questions into the large language model based on knowledge filtering to obtain corresponding common sense knowledge, so as to construct first question-knowledge pairs; a data classification module configured to classify the first question-knowledge pairs by evaluating the influence degree of the corresponding common sense knowledge on the corresponding answers of the plurality of questions and generate corresponding level labels, to obtain second question-knowledge pair data including the level labels; a reward value determination module configured to input the second question-knowledge pair data and the data set into a preset reward model to obtain a corresponding reward value output; and a parameter adjustment module configured to adjust each parameter of the preset reward model according to the reward value output to obtain a trained reward model, so as to realize the inference enhancement of the large language model based on knowledge filtering.
[0013] Optionally, the data acquisition module may further be configured to acquire a question to be inferred related to common sense knowledge, and the inference enhancement device may further include an inference module, wherein the inference module is configured to input the question to be inferred into the large language model based on knowledge filtering to obtain corresponding multiple inference chains; and input the corresponding multiple inference chains into the trained reward model to obtain valid inference chains corresponding to the question to be inferred as inference chains enhanced by inference.
[0014] Optionally, the inference module may further be configured to input the question to be inferred and the inference chains enhanced by inference into the large language model based on knowledge filtering to obtain multiple answers to the question to be inferred after multiple inferences, and use the answer with the highest marginal probability among the multiple answers as the final answer to the question to be inferred to obtain an inference answer enhanced by inference.
[0015] Optionally, the operation of the data classification module to classify the first question-knowledge pair and generate a corresponding level label by evaluating the influence degree of the corresponding common sense knowledge on the corresponding answers to multiple questions, and obtain the second question-knowledge pair data including the level label may include: inputting the data set and the first question-knowledge pair into the large language model based on knowledge filtering to obtain a first answer and a second answer, where the first answer is obtained when the large language model based on knowledge filtering uses the corresponding common sense knowledge to solve the answer, and the second answer is obtained when the large language model based on knowledge filtering does not use the corresponding common sense knowledge to solve the answer; evaluating the influence degree of the corresponding common sense knowledge on the corresponding answers to multiple questions according to the first answer and the second answer, classifying the first question-knowledge pair and generating a corresponding level label, and obtaining the second question-knowledge pair data including the level label.
[0016] Optionally, the operation of the parameter adjustment module to adjust each parameter of the preset reward model according to the reward value output may include: calculating a loss function value for adjusting each parameter through the following formula according to the reward value output:
[0017] L(θ) = -ylog(f(q,k;θ)) - (1 - y)log(1 - f(q,k;θ)),
[0018] where θ represents the model parameters of the preset reward model, L represents the loss function, y represents the level label, f represents the reward value predicted by the preset reward model, q represents the question, and k represents the corresponding common sense knowledge; adjusting each parameter based on the loss function value to obtain the trained reward model.
[0019] In another general aspect, a computer program product is provided, where the computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the inference enhancement method of the large language model based on knowledge filtering as described above is implemented.
[0020] In another general aspect, a computer-readable storage medium is provided, and when the instructions in the computer-readable storage medium are executed by a processor of an electronic device / server, the electronic device / server can execute the inference enhancement method of the large language model based on knowledge filtering as described above.
[0021] In another general aspect, a computing device is provided, the computing device including: at least one processor; at least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the method for enhancing inference of a large language model based on knowledge filtering as described above.
[0022] According to the method and apparatus for enhancing common sense reasoning of a large language model based on knowledge filtering according to an embodiment of the present disclosure, by utilizing knowledge filtering, the common sense reasoning ability of the large language model is enhanced, and the quality of the thought chain generated by the large language model and the accuracy of answering common sense reasoning tasks are greatly improved. In addition, compared with other existing methods, the method for enhancing common sense reasoning of a large language model based on knowledge filtering proposed by the present disclosure maintains a good balance between the positive effects and the adverse effects brought, and can effectively avoid damaging the original reasoning ability of the large language model while further improving the positive effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Through the following description with reference to the drawings showing embodiments, the above and other objects and features of the embodiments of the present disclosure will become clearer, wherein:
[0024] Figure 1 is a flowchart showing a method for enhancing inference of a large language model based on knowledge filtering according to an embodiment of the present disclosure;
[0025] Figure 2 is a schematic diagram showing a system architecture for applying the above inference enhancement method according to an embodiment of the present disclosure;
[0026] Figure 3 is a schematic diagram showing an existing common sense reasoning task;
[0027] Figure 4 is a schematic diagram showing a marginal consistency reasoning method according to an embodiment of the present disclosure;
[0028] Figure 5 is a block diagram showing an apparatus for enhancing inference of a large language model based on knowledge filtering according to an embodiment of the present disclosure;
[0029] Figure 6 is a block diagram showing a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The following specific embodiments are provided to assist the reader in obtaining a comprehensive understanding of the methods, devices, and / or systems described herein. However, after understanding the disclosure of this application, various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will be apparent. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of this application, except for operations that must occur in a specific order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.
[0031] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein like reference numerals always refer to like elements. The embodiments will be described below with reference to the accompanying drawings in order to explain the present disclosure.
[0032] As described above, the common sense reasoning ability of large models can be further improved by two ways: retrieval enhancement and self-enhancement.
[0033] Regarding the retrieval enhancement method, as Figure 3 shown in Example 1 of [], this method retrieves knowledge related to the question from the knowledge graph and integrates it into the input of the model as supplementary information. In this case, due to the limited coverage of common sense knowledge in the knowledge graph and the fact that the retriever can only capture semantic similarities between entities, it is difficult to recall effective information in complex common sense reasoning scenarios (such as event-based reasoning). Referring again to Figure 3 Example 1 of [], in the WinoGrande question, the model needs common sense knowledge to describe the relationship between "becoming a better surgeon" and "getting easier cases", and the most relevant knowledge in the common sense knowledge graph (e.g., ATOMIC-2020), "(X person gets stitches, y effect, Y person will gain more medical experience)", is still far from the actual needs.
[0034] Regarding the self-enhancement method, as Figure 3 shown in Examples 2 and 3 of [], this method uses the Chain of Thought (CoT) technology to endow large models with the ability to generate the knowledge required for reasoning. Existing CoT-based methods have the defects of noisy knowledge and invalid reasoning. For example, for noisy knowledge, the reasoning process generated by the large model itself may contain serious noise, which will have a negative impact on reasoning. For example, in Figure 3 Example 2 of [], the generated knowledge, "To get up late means people tend to go to bed late and get up late", belongs to noisy information, which will cause the large model to give a wrong answer. In addition, for invalid reasoning, there are cases where even if reasonable knowledge is provided to the large model, it may still lead to wrong answers, which can be expressed as the "invalid reasoning" problem. For example, in Figure 3In Example 3, although the reasoning path "the tomb is not large enough to fully accommodate the body" is correct for this problem, the large model still fails to draw the correct conclusion based on this.
[0035] To solve the above and / or other problems, the present disclosure proposes a method and device for enhancing the reasoning of a large language model based on knowledge filtering, so as to enhance the common sense reasoning ability of the large model by using effective knowledge. By using knowledge filtering to enhance the common sense reasoning ability of the large language model, the quality of the thinking chain generated by the large language model and the accuracy of answering common sense reasoning tasks are greatly improved.
[0036] The following refers to Figures 1 to 6 and describes in detail a method and device for enhancing common sense reasoning of a large language model based on knowledge filtering according to an embodiment of the present disclosure.
[0037] Figure 1 is a flowchart showing a method 100 for enhancing the reasoning of a large language model based on knowledge filtering according to an embodiment of the present disclosure. Figure 2 is a schematic diagram of a system architecture applying the above reasoning enhancement method according to an embodiment of the present disclosure.
[0038] Referring to Figure 1 , according to an embodiment of the present disclosure, in step S101, a data set including a plurality of questions for common sense knowledge and corresponding a plurality of answers is obtained.
[0039] In the present disclosure, the data in the data set is in text form. For example, Figure 2 and Figure 3 the question and answer of the text type shown in.
[0040] According to an embodiment of the present disclosure, in step S102, the plurality of questions are input into a large language model based on knowledge filtering to obtain corresponding common sense knowledge, so as to construct a first question-knowledge pair.
[0041] According to an embodiment of the present disclosure, in step S103, by evaluating the influence degree of the corresponding common sense knowledge on the corresponding answers of the plurality of questions, the first question-knowledge pair is classified and a corresponding level label is generated to obtain a second question-knowledge pair data including the level label.
[0042] As an example, step S103 may further include steps S1031 and S1032:
[0043] In step S1031, the data set and the first question-knowledge pair are input into a large language model based on knowledge filtering to obtain a first answer and a second answer.
[0044] Here, the first answer is obtained when the large language model based on knowledge filtering uses the corresponding common sense knowledge to solve the answer, and the second answer is obtained when the large language model based on knowledge filtering does not use the corresponding common sense knowledge to solve the answer.
[0045] In step S1032, according to the first answer and the second answer, the influence degree of the corresponding common sense knowledge on the corresponding answers to multiple questions is evaluated, the first question-knowledge pair is classified, and the corresponding level label is generated to obtain the second question-knowledge pair data including the level label.
[0046] That is to say, according to the above embodiments, by evaluating the common sense knowledge related to the questions generated by the large model, these question-knowledge pairs are classified to construct training data. Further, since the large model itself contains a large amount of common sense knowledge, the large model itself can be used as a knowledge source. For the questions in the training data, first, a context learning prompt model can be used to generate multiple pieces of relevant knowledge; then, the large model is used to answer the question in two cases of accessing and not accessing this knowledge respectively, and the knowledge is rated according to its correctness. For example, for a question-knowledge pair, the effectiveness of the knowledge in helping the model answer the question correctly decreases as the level is marked from 0 to 3. That is, level 0 knowledge (knowledge with a level label of 0) can enhance the ability of the large model and help the large model answer questions that are originally difficult to answer. On the contrary, level 3 knowledge (knowledge with a level label of 3) will cause the large model to generate wrong answers to common sense questions that it can usually answer correctly. Therefore, the above knowledge level labels including, for example, four levels can be used to evaluate the effectiveness and / or harmfulness of knowledge to provide a supervised learning signal for training the reward model.
[0047] According to an embodiment of the present disclosure, for each question-knowledge pair, it is classified into samples with different level labels according to the reward value of the knowledge, and trained using the objective of contrast learning, so that the reward model can correctly identify valid knowledge and noise knowledge.
[0048] According to an embodiment of the present disclosure, in step S104, the second question-knowledge pair data and the data set are input into a preset reward model to obtain a corresponding reward value output.
[0049] According to an embodiment of the present disclosure, in step S105, according to the reward value output, each parameter of the preset reward model is adjusted to obtain a trained reward model, so as to realize the inference enhancement of the large language model based on knowledge filtering.
[0050] As an example, step S105 may further include steps S1051 and S1052:
[0051] In step S1051, according to the output of the reward value, the loss function value for adjusting each parameter is calculated by the following formula (1):
[0052] L(θ)=-y log(f(q,k;θ))-(1-y)log(1-f(q,k;θ)) (1)
[0053] where θ represents the model parameters of the preset reward model, L represents the loss function (e.g., as shown by the "BCE loss" in Figure 2 ), y represents the level label, f represents the reward value predicted by the preset reward model, q represents the question, and k represents the corresponding common sense knowledge; each parameter is adjusted based on the loss function value to obtain the trained reward model.
[0054] As an example, according to the above embodiments, by using contrastive learning to train the reward model on the training data, the reward model can distinguish between noisy knowledge and high-quality knowledge. Specifically, after collecting a set of data including question and knowledge pairs and their corresponding knowledge levels as described above, for the purpose of training efficiency, these data can be further classified into positive sample data (hereinafter referred to as positive samples) and negative sample data (hereinafter referred to as negative samples), and corresponding labels are further assigned. For example, considering the contribution of knowledge in answering questions, if the level of a piece of knowledge is 0 or 1, it is defined as a positive sample of the question, otherwise, if the level of a piece of knowledge is 2 or 3, it is defined as a negative sample. For example, in the actual execution process, questions that are only associated with positive samples or only associated with negative samples can be deleted because such questions do not contribute much to the optimization of the model.
[0055] Specifically, as an example, first, for each question q in the training data, the context learning prompt model can be used to generate multiple relevant knowledge K q . For example, for each question q, the generated knowledge K q contains multiple knowledge fragments k. When predicting the answer to a question, the large model needs to consider the following two cases: Case 1, generating an answer without knowledge, that is, directly using the question q to generate an answer r(q), and its expression is: r(q)=M(q,Pd), where Pd is the prompt for generating a direct answer; Case 2, generating an answer with knowledge, that is, generating an answer r(q,k) in the presence of the knowledge fragment k, and its expression is: r(q,k)=M(q,Pk,k), where Pk is the prompt for generating an answer based on knowledge.
[0056] Then, according to the correctness of the generated answers r(q) and r(q,k), the knowledge fragment k can be divided into four confidence levels (level labels) from 0 to 4, and the specific description is as follows (hereinafter, a represents the correct answer):
[0057] (1) Useful (e.g., Figure 2 Level 0 as shown): Corresponding to the case where r(q) ≠ a and r(q,k) = a, and indicating that the knowledge fragment k helps to correctly answer the question.
[0058] (2) Harmless (e.g., Figure 2 Level 1 as shown): Corresponding to the case where r(q) = a and r(q,k) = a, and indicating that the knowledge fragment k does not have a negative impact on the answer.
[0059] (3) Useless (e.g., Figure 2 Level 2 as shown): Corresponding to the case where r(q) ≠ a and r(q,k) ≠ a, and indicating that the knowledge fragment k is ineffective for the answer.
[0060] (4) Harmful (e.g., Figure 2 Level 3 as shown): Corresponding to the case where r(q) = a and r(q,k) ≠ a, and indicating that the knowledge fragment k leads to an incorrect answer.
[0061] That is to say, the effectiveness of knowledge decreases from level 0 to 3, while the harmfulness increases from level 0 to 3. These levels provide the supervision signals for training the reward model.
[0062] In addition, for the training objective (i.e., filtering out noisy knowledge), the present disclosure enables the reward model to assign a higher score (i.e., a high reward value) to effective knowledge and a lower score to noisy knowledge. For example, a reward model based on DeBERTa can be used as a cross-encoder to encode both the question and the knowledge simultaneously, and then generate a confidence score (i.e., the reward value) between 0 and 1.
[0063] Specifically, for each question q in the training data, multiple relevant knowledge Ks q as described above and the knowledge fragment k, first, obtain a set of (q,k) data pairs and their corresponding knowledge levels; then, according to the level of knowledge, divide the data into positive samples (levels 0 and 1) and negative samples (levels 2 and 3); finally, by scoring the (q,k) pairs, train the reward model to optimize its ability to identify effective knowledge.
[0064] By adopting a reward model including the above loss function, defining the confidence of knowledge according to the contribution of knowledge to question answering, and using it as the supervision signal for training the reward model, effective filtering of the noisy knowledge generated by the large language model is achieved.
[0065] Additionally, the inference enhancement method 100 for large language models based on knowledge filtering may further include: obtaining an inference question related to common sense knowledge; inputting the inference question into the large language model based on knowledge filtering to obtain corresponding multiple inference chains (or called chains of thought); inputting the corresponding multiple inference chains into the trained reward model to obtain an effective inference chain corresponding to the inference question, which is used as the inference chain after inference enhancement.
[0066] According to an embodiment of the present disclosure, the above-mentioned large model common sense inference enhancement method based on knowledge filtering trains a reward model by using the influence of knowledge on answering questions as a supervision signal, and filters the generated inference chains according to the reward values of the reward model to obtain an inference chain after inference enhancement for subsequent obtaining of the most accurate inference answer.
[0067] Furthermore, the inference enhancement method 100 for large language models based on knowledge filtering may further include: inputting the inference question and the inference chain after inference enhancement into the large language model based on knowledge filtering to obtain multiple answers to the inference question after multiple inferences, and taking the answer with the highest marginal probability among the multiple answers as the final answer to the inference question, thus obtaining the inference answer after inference enhancement.
[0068] As an example, according to the above embodiment, when using a large model to implement downstream inference, first use the large model to generate multiple inference chains, and use the reward model to screen out the noise knowledge (filter noise knowledge) among them. Combine the remaining multiple high-quality inference chains, and use the large model to answer multiple rounds based on this knowledge, enabling the large model to reason repeatedly, thereby reducing the interference of these low-quality contents on the final inference result.
[0069] Specifically, for each question, use the reward model to score the generated knowledge, select several pieces of knowledge with the highest scores, and connect them into a valid inference path; then, the valid inference path can be integrated into the input, and the large model is prompted to perform multiple rounds of inference; finally, output the result determined by majority voting on the answers. Through the above inference process, the uncertainty in the sampling process of the large model can be reduced, thereby improving the ability of the large model in common sense inference tasks.
[0070] For the marginal consistency inference process described above, by performing multiple rounds of inference using an effective inference path and selecting the answer with the highest marginal probability, the most accurate answer can be obtained. The corresponding formula of its principle is as follows:
[0071] argmax a′ P(a′|q)≈argmax a′ P(a′|k * ,q) (2)
[0072] P(a′|k* , q) ≈ frequency(a′) / n ∝ frequency(a′) (3)
[0073] where k * represents an effective reasoning chain, n represents the number of samplings, q represents the question, and a′ represents the answer. From the above formulas (2) and (3), it can be obtained that by multi-sampling to select the answer with the highest frequency, that is, the most likely correct target. Through this method of filtering and multi-sampling voting, the stability of reasoning can be improved.
[0074] According to an embodiment of the present disclosure, by performing multi-round reasoning for each question using an effective reasoning path and selecting the answer with the highest marginal probability as the final answer, invalid reasoning can be effectively reduced. In addition, compared with traditional Chain of Thought (CoT) - like methods, the problem that traditional methods may generate incorrect outputs when the probability distribution of candidate answers is relatively uniform is solved.
[0075] According to an embodiment of the present disclosure, by using the influence of knowledge on answering questions as a supervision signal to train a reward model, and filtering the generated reasoning chains according to the reward values of the reward model, the model repeatedly thinks on high-quality reasoning chains to generate the final answer, thereby enhancing the common sense reasoning level of the large language model.
[0076] Next, with reference to Figure 2 to illustrate the system architecture of an inference enhancement method for applying a large language model based on knowledge filtering according to an embodiment of the present disclosure.
[0077] Figure 2 is a schematic diagram showing the system architecture of applying the above inference enhancement method according to an embodiment of the present disclosure.
[0078] As Figure 2 shown, the system architecture of applying the above inference enhancement method according to the present disclosure may include three parts: a knowledge pooling processing stage, a reward model training processing stage, and a marginal consistency reasoning processing stage, which are specifically as follows:
[0079] In the knowledge pooling processing stage, the large model is run with and without knowledge provided respectively, the performance of different knowledge is recorded, and then the effectiveness of the knowledge is rated based on correctness to construct training data; in the reward model training processing stage, for each question-knowledge pair, it is divided into positive and negative samples according to the score of the knowledge, and trained using the objective of contrastive learning, so that the reward model can correctly identify valid and noisy knowledge; in the marginal consistency reasoning processing stage, during downstream reasoning, high-quality generated thought chains are screened out by the reward model and connected into valid thought chains. Based on this valid thought chain, the large model repeatedly thinks, samples multiple correct answers and obtains the final reasoning answer through majority voting as the final answer to the question.
[0080] According to the present disclosure, in order to verify the effectiveness of the above-mentioned method for enhancing common sense reasoning of large language models based on knowledge filtering, comprehensive experimental verification of the above method was carried out on four complex common sense reasoning benchmarks, and the specific results can be seen in Table 1 below.
[0081] In Table 1, the test corpus includes the following datasets: WinoGrande, HellaSwag, SocialIQA, PIQA. The methods for comparison include four categories, namely, the Few-shot method, the fine-tuning method (including Roberta-large, Unified QA, Unicorn), the retrieval enhancement method (including BM25+LLM, DPR+LLM), and the self-enhancement method (including CoT, CoT-SC, Self-Refine, Least-to-Most). The above method of the present disclosure is denoted as LINKED.
[0082] Table 1
[0083]
[0084] In Table 1, ACC represents the accuracy rate of the answer result, and EPS represents the effect preservation score and can measure both the positive and negative impacts of the knowledge enhancement method at the same time. The specific calculation formula of EPS is as follows:
[0085]
[0086] Among them, ES represents the effectiveness score, and the meaning of its numerator is as follows: the number of questions that were answered incorrectly without adding knowledge k but were answered correctly after adding k, a * represents the correct answer, and r represents the model output; the meaning of its denominator is as follows: the number of all questions that were answered incorrectly (without adding knowledge).
[0087] Among them, PS represents the preservation score, and the meaning of its numerator is as follows: the number of questions that were answered correctly without adding knowledge k However, the number of questions with incorrect answers after adding k, a * represents the correct answer, and r represents the model output; the denominator has the following meaning: the number of questions that are directly answered (without adding knowledge) correctly.
[0088] From the experimental results in Table 1, it can be seen that according to the above method LINKED of the present disclosure, it is superior to the current state-of-the-art baseline (the data underlined in Table 1), and the accuracy rate has increased by up to 9.0%. For example, on the WinoGrande dataset, the accuracy rate of the method LINKED of the present disclosure has increased by 9.0%.
[0089] In addition, in order to measure the positive and negative impacts of injecting knowledge, the present disclosure proposes a new metric EPS. The method LINKED of the present disclosure has increased EPS by an average of 5.4%, which indicates that the method of the present disclosure effectively avoids damaging the original reasoning ability of the large model while introducing effective knowledge. Therefore, the method of the present disclosure maintains a good balance between effectiveness and harmfulness.
[0090] The results of experiments on multiple datasets show that the method of the present disclosure significantly improves the quality of the thought chain generated by the large model and the accuracy rate of answering common sense reasoning tasks.
[0091] Next, with reference to Figure 5 Describe an inference enhancement device for a large language model based on knowledge filtering according to an embodiment of the present disclosure. Figure 5 FIG. is a block diagram showing an inference enhancement device 500 for a large language model based on knowledge filtering according to an embodiment of the present disclosure.
[0092] With reference to Figure 5 , the inference enhancement device 500 for a large language model based on knowledge filtering according to an embodiment of the present disclosure may include a data acquisition module 510, a data construction module 520, a data classification module 530, a reward value determination module 540, and a parameter adjustment module 550.
[0093] According to an embodiment of the present disclosure, the data acquisition module 510 may execute: acquiring a data set including a plurality of questions and corresponding answers for common sense knowledge.
[0094] According to an embodiment of the present disclosure, the data construction module 520 may execute: inputting a plurality of questions into a large language model based on knowledge filtering to obtain corresponding common sense knowledge, so as to construct a first question-knowledge pair.
[0095] According to an embodiment of the present disclosure, the data classification module 530 may execute: classifying the first question-knowledge pair by evaluating the influence degree of the corresponding common sense knowledge on the corresponding answers of a plurality of questions and generating a corresponding level label, so as to obtain a second question-knowledge pair data including the level label.
[0096] As an example, the data classification module 530 classifies the first question-knowledge pair by evaluating the influence degree of the corresponding common sense knowledge on the corresponding answers to multiple questions and generates a corresponding level label. The operations for obtaining the second question-knowledge pair data including the level label may include the following operations 1) and 2):
[0097] In operation 1), the data set and the first question-knowledge pair are input into the large language model based on knowledge filtering to obtain a first answer and a second answer.
[0098] Here, the first answer is obtained when the large language model based on knowledge filtering uses the corresponding common sense knowledge to solve the answer, and the second answer is obtained when the large language model based on knowledge filtering does not use the corresponding common sense knowledge to solve the answer.
[0099] In operation 2), according to the first answer and the second answer, the influence degree of the corresponding common sense knowledge on the corresponding answers to multiple questions is evaluated, the first question-knowledge pair is classified and a corresponding level label is generated to obtain the second question-knowledge pair data including the level label.
[0100] According to an embodiment of the present disclosure, the reward value determination module 540 may execute: inputting the second question-knowledge pair data and the data set into a preset reward model to obtain a corresponding reward value output.
[0101] According to an embodiment of the present disclosure, the parameter adjustment module 550 may execute: adjusting each parameter of the preset reward model according to the reward value output to obtain a trained reward model, so as to realize the inference enhancement of the large language model based on knowledge filtering.
[0102] As an example, the operations for the parameter adjustment module 550 to adjust each parameter of the preset reward model according to the reward value output to obtain a trained reward model may include: calculating the loss function value for adjusting each parameter through the above formula (1) according to the reward value output; adjusting each parameter based on the loss function value to obtain a trained reward model.
[0103] The inference enhancement device 500 of the large language model based on knowledge filtering according to an embodiment of the present disclosure may further include an inference module. As an example, after the inference module obtains a question to be inferred related to common sense knowledge via the data acquisition module 510, the inference module may execute: inputting the question to be inferred into the large language model based on knowledge filtering to obtain corresponding multiple inference chains; inputting the corresponding multiple inference chains into the trained reward model to obtain a valid inference chain corresponding to the question to be inferred as an inference chain after inference enhancement.
[0104] In addition, the inference module may also perform the following: input the problem to be inferred and the inference chain enhanced through inference into a large language model based on knowledge filtering to obtain multiple answers to the problem to be inferred after multiple inferences, and use the answer with the highest marginal probability among the multiple answers as the final answer to the problem to be inferred, thereby obtaining the inference answer enhanced through inference.
[0105] It should be noted that the operations performed by the above respective structural blocks may be similar to the related content described with reference to Figure 1 and will not be elaborated here.
[0106] Figure 6 FIG. is a block diagram showing a computing device 600 according to an embodiment of the present disclosure.
[0107] With reference to Figure 6 , the computing device 600 according to an embodiment of the present disclosure may include a processor 610 and a memory 620. The processor 610 may include (but is not limited to) a central processing unit (CPU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a system on chip (SoC), a microprocessor, an application specific integrated circuit (ASIC), etc. The memory 620 may store computer executable instructions to be executed by the processor 610. The memory 620 includes high-speed random access memory and / or non-volatile computer-readable storage media. When the processor 610 executes the computer executable instructions stored in the memory 620, the inference enhancement method of the large language model based on knowledge filtering as described above can be implemented.
[0108] The inference enhancement method of the large language model based on knowledge filtering according to an embodiment of the present disclosure can be written as computer programs / instructions to form a computer program product and stored on a computer-readable storage medium. When the computer programs / instructions are executed by a processor, the inference enhancement method of the large language model based on knowledge filtering as described above can be implemented. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device / server, the electronic device / server is enabled to execute the inference enhancement method of the large language model based on knowledge filtering as described above. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer such that the processor or computer can execute the computer program. In one example, the computer program and any associated data, data files and data structures are distributed across a networked computer system such that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.
[0109] The common sense inference enhancement method and device of the large language model based on knowledge filtering according to an embodiment of the present disclosure enhance the common sense inference ability of the large language model by using knowledge filtering, greatly improving the quality of the thinking chain generated by the large language model and the accuracy of answering common sense inference tasks.
[0110] In addition, compared with other existing methods, the common sense inference enhancement method of the large language model based on knowledge filtering proposed by the present disclosure maintains a good balance between the positive effects and adverse effects brought, and can effectively avoid damaging the original inference ability of the large language model while further enhancing the positive effects.
[0111] In addition, the method and device for enhancing commonsense reasoning of large language models based on knowledge filtering proposed by the present disclosure significantly improve the quality of the thought chain generated by the large model and the accuracy of answering commonsense reasoning tasks. In addition, the method and device for enhancing commonsense reasoning of large language models based on knowledge filtering proposed by the present disclosure maintain a good balance between effectiveness and harmfulness.
[0112] Although some embodiments of the present disclosure have been disclosed and described, those skilled in the art should understand that these embodiments can be modified and varied without departing from the concept and spirit of the present disclosure as defined by the claims and their equivalents.
Claims
1. A method for enhancing reasoning of a large language model based on knowledge filtering, characterized in that: The reasoning enhancement method comprises: Obtain a dataset including a plurality of questions for common sense knowledge and a plurality of corresponding answers; Inputting the multiple questions into a large language model based on knowledge filtering to obtain corresponding common sense knowledge to construct a first question-knowledge pair; By evaluating the influence of the corresponding common sense knowledge on the corresponding answers to the multiple questions, the first question-knowledge pair is classified and the corresponding level label is generated, thereby obtaining the second question-knowledge pair data containing the level label; Inputting the second question-knowledge pair data and the data set into a preset reward model to obtain a corresponding reward value output; According to the reward value output, various parameters of the preset reward model are adjusted to obtain a trained reward model to achieve reasoning enhancement of the large language model based on knowledge filtering.
2. The reasoning enhancement method according to claim 1, characterized in that: The reasoning enhancement method further includes: Get questions to be reasoned about that are related to common sense knowledge; Inputting the problem to be inferred into the large language model based on knowledge filtering to obtain corresponding multiple reasoning chains; The corresponding multiple reasoning chains are input into the trained reward model to obtain valid reasoning chains corresponding to the problem to be reasoned, as reasoning chains enhanced by reasoning.
3. The reasoning enhancement method according to claim 2, characterized in that: The reasoning enhancement method further includes: The problem to be inferred and the reasoning chain enhanced by reasoning are input into the large language model based on knowledge filtering, and multiple answers to the problem to be inferred after multiple reasonings are obtained, and the answer with the highest marginal probability among the multiple answers is used as the final answer to the problem to be inferred, so as to obtain the reasoning answer enhanced by reasoning.
4. The reasoning enhancement method according to claim 1, characterized in that: The step of classifying the first question-knowledge pair and generating corresponding level labels by evaluating the influence of the corresponding common sense knowledge on the corresponding answers to the multiple questions, and obtaining the second question-knowledge pair data containing the level labels comprises: Inputting the data set and the first question-knowledge pair into the large language model based on knowledge filtering to obtain a first answer and a second answer, wherein the first answer is obtained when the large language model based on knowledge filtering uses the corresponding common sense knowledge to solve the answer, and the second answer is obtained when the large language model based on knowledge filtering does not use the corresponding common sense knowledge to solve the answer; According to the first answer and the second answer, the influence of the corresponding common sense knowledge on the corresponding answers to multiple questions is evaluated, the first question-knowledge pairs are classified and corresponding level labels are generated, and the second question-knowledge pair data containing the level labels are obtained.
5. The reasoning enhancement method according to claim 1, characterized in that: The step of adjusting various parameters of the preset reward model according to the reward value output to obtain a trained reward model comprises: According to the reward value output, the loss function value used to adjust the various parameters is calculated by the following formula: L(θ)=-ylog(f(q,k;θ))-(1-y)log(1-f(q,k;θ)) Among them, θ represents the model parameters of the preset reward model, L represents the loss function, y represents the level label, f represents the reward value predicted by the preset reward model, q represents the question, and k represents the corresponding common sense knowledge; The various parameters are adjusted based on the loss function value to obtain a trained reward model.
6. A reasoning enhancement device for a large language model based on knowledge filtering, characterized in that: The reasoning enhancement device comprises: A data acquisition module is configured to: acquire a data set including a plurality of questions for common sense knowledge and a plurality of corresponding answers; The data construction module is configured to: input the plurality of questions into a large language model based on knowledge filtering to obtain corresponding common sense knowledge to construct a first question-knowledge pair; A data classification module is configured to: classify the first question-knowledge pair and generate corresponding level labels by evaluating the influence of the corresponding common sense knowledge on the corresponding answers to the multiple questions, and obtain second question-knowledge pair data containing the level labels; A reward value determination module is configured to: input the second question-knowledge pair data and the data set into a preset reward model to obtain a corresponding reward value output; The parameter adjustment module is configured to: adjust various parameters of the preset reward model according to the reward value output to obtain a trained reward model to achieve reasoning enhancement of the large language model based on knowledge filtering.
7. The reasoning enhancement device according to claim 6, characterized in that: The data acquisition module is further configured to: acquire questions to be reasoned related to common sense knowledge, and the reasoning enhancement device further includes a reasoning module, Wherein, the reasoning module is configured as follows: Inputting the problem to be inferred into the large language model based on knowledge filtering to obtain corresponding multiple reasoning chains; The corresponding multiple reasoning chains are input into the trained reward model to obtain valid reasoning chains corresponding to the problem to be reasoned, as reasoning chains enhanced by reasoning.
8. The reasoning enhancement device according to claim 7, characterized in that: The reasoning module is further configured to: The problem to be inferred and the reasoning chain enhanced by reasoning are input into the large language model based on knowledge filtering, and multiple answers to the problem to be inferred after multiple reasonings are obtained, and the answer with the highest marginal probability among the multiple answers is used as the final answer to the problem to be inferred, so as to obtain the reasoning answer enhanced by reasoning.
9. A computer program product, characterized in that The computer program product comprises a computer program / instruction, and when the computer program / instruction is executed by a processor, the method for enhancing reasoning of a large language model based on knowledge filtering according to any one of claims 1 to 5 is implemented.
10. A computing device, characterized in that The computing device includes: at least one processor; and at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, cause the at least one processor to execute the reasoning enhancement method for a large language model based on knowledge filtering as described in any one of claims 1 to 5.