Iterative background knowledge extraction method and system based on large language model, and medium
By adopting an iterative background knowledge extraction method based on large language models in the textbook Q&A system, the problem of difficulty in accurately extracting related background knowledge in long text courses is solved, and more efficient and accurate knowledge extraction is achieved, which is suitable for deep semantic understanding of multimodal content.
Patent Information
- Application Number
- CN202411991455.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to accurately extract background knowledge directly related to specific issues when dealing with long text courses, especially inadequate performance in multimodal content processing and deep semantic understanding.
The iterative background knowledge extraction method based on the large language model is adopted. By finding the paragraphs most relevant to the problem in the course text, combining the word frequency-inverse document frequency (TF-IDF) method and evidence theory, the knowledge extraction process is iteratively optimized until the reasoning result of the large language model is consistent with the answer with the highest score of the total belief function.
It improves the accuracy and efficiency of textbook question-and-answer system in long text processing, can more accurately locate and extract knowledge related to the problem, adapt to the processing of multimodal content, and provide a deeper semantic understanding.
Smart Images

Figure CN120012899A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of computer semantic analysis, and specifically relates to an iterative background knowledge extraction method, system and medium based on a large language model. Background Art
[0002] Existing research works often use text similarity or search engines to obtain background knowledge related to the question and retain several sentences with high scores when dealing with long text courses in textbook question answering tasks. The background technology of iterative background knowledge extraction methods based on large language models covers a range of technologies and methods for enhancing the efficiency and accuracy of textbook question answering systems when dealing with long course materials. The background of this method is built on the limitations of traditional question answering systems, which usually rely on simple search algorithms or basic machine learning techniques such as support vector machines and decision trees. These traditional methods often show obvious shortcomings when dealing with complex queries and understanding deep textual relationships.
[0003] With the emergence of pre-trained language models such as BERT and GPT, textbook question answering systems have seen significant performance improvements. These large language models are able to learn and capture complex patterns and deep relationships in language by pre-training on large-scale corpora. By applying these models to textbook question answering tasks, researchers can significantly improve the system's ability to understand text, especially when dealing with long texts and complex semantics.
[0004] Despite the progress brought by large language models, directly applying these models to long text processing remains challenging, especially in accurately extracting knowledge fragments that are directly relevant to a specific question. Summary of the invention
[0005] The purpose of the present invention is to address the above-mentioned problems in the prior art and to provide an iterative background knowledge extraction method, system and medium based on a large language model, which can effectively and accurately extract knowledge fragments directly related to specific questions from long textbook texts, adapt to processing multimodal content, and effectively combine different forms of information to provide accurate answers.
[0006] In order to achieve the above object, the present invention has the following technical solutions:
[0007] In a first aspect, a method for iterative background knowledge extraction based on a large language model is provided, comprising:
[0008] Find several paragraphs in the course text that are most relevant to the questions and options as candidate knowledge sets, and combine the candidate knowledge sets with the corresponding relationship descriptions to form a background knowledge candidate set;
[0009] For a sentence s in the background knowledge candidate set, the question q and the candidate answer a are converted into a declarative proposition h through the large language model, and the confidence score of the sentence s and the candidate answer a of the question q is obtained;
[0010] Use evidence theory to synthesize the confidence scores of all candidate answers a to obtain the total belief function score of proposition h, and determine the answer with the highest total belief function score;
[0011] Determine whether the answer with the highest total belief function score is consistent with the answer inferred by the large language model. If not, modify the confidence score of the large language model for candidate answer a or modify the answer inferred. Iterate until the answer inferred by the large language model is consistent with the answer with the highest total belief function score, and output the answer inferred by the large language model.
[0012] As a preferred solution, the step of finding several paragraphs in the course text that are most relevant to the questions and options as candidate knowledge sets is implemented using the term frequency-inverse document frequency TF-IDF method. The importance of vocabulary is measured by the inverse document frequency, combined with the repetition frequency of stop words StopWords and topic vocabulary in different text paragraphs. When a word appears multiple times in multiple paragraphs, the TF-IDF value decreases, indicating that the corresponding word has a low correlation with the specific question or option.
[0013] As a preferred solution, the step of searching for several paragraphs in the course text that are most relevant to the questions and options as candidate knowledge sets, and splicing the candidate knowledge sets with the corresponding relationship descriptions into a background knowledge candidate set includes:
[0014] For a textbook question-answering problem q∈Q, concatenate the question and the corresponding answer candidates, and express the concatenated result as Each item in the set represents a word, n q is the total number of words obtained after concatenation;
[0015] For the paragraph set C in the problem background course q ={p1, p2, ..., p j}, each paragraph p in the collection j ∈C q It is regarded as an independent document. The score of each word ti is calculated by the term frequency-inverse document frequency TF-IDF method. The sum is used to get the score of the corresponding paragraph for the question option. The calculation expression is as follows:
[0016]
[0017] In the formula, Score jq For paragraph pj Relevance score for question q;
[0018] TF-IDF t,d =TF t,d ×IDF t
[0019] Where TFt,d is the term frequency, which represents the number of times term t appears in document d; IDF t is the inverse document frequency, which represents the rarity of term t in all documents. The calculation expression is as follows:
[0020]
[0021] Where N is the total number of documents, n (t) is the number of documents containing t;
[0022] Sort by score, retain a number of paragraphs, and concatenate them with the relationship description to form a candidate set of background knowledge s q is the number of sentences in the background knowledge candidate set for question q.
[0023] As a preferred solution, in the step of obtaining the confidence score of a sentence s and a candidate answer a to question q, the confidence score is performed from three dimensions: relevance, support, and sufficiency; wherein relevance represents the degree of connection between background knowledge and the proposition; support represents whether the background knowledge can support the corresponding answer; sufficiency represents whether the background knowledge provides enough information to fully support the answer, including all necessary details and conditions;
[0024] The support score range is set to [-1, 1], the relevance and sufficiency score range is set to (0, 1]; the confidence score is recorded as E sqa , the support, relevance and sufficiency scores are denoted as e s , e r , e a ,but:
[0025] E sqa =max(e s ×e r ×e a ,0)
[0026] In the formula, the confidence score E sqa is the belief function value of sentence s for the proposition h consisting of question q and candidate answer a.
[0027] As a preferred solution, the confidence scores of all candidate answers a are synthesized using evidence theory to obtain the total belief function score of proposition h, and the answer with the highest total belief function score is determined, including:
[0028] Evidence theory defines the background knowledge candidate set K q As a set of candidate evidence sources, the belief function score dictionary M of the candidate evidence source set on the proposition set H v As input, for a particular source of evidence k and a proposition h consisting of q and a, there exists:
[0029] M v [k][h]=E sqa
[0030] The output of the evidence theory is the total belief function score Mc[h] for each proposition h∈H; after concatenating the total belief function scores of all propositions, the softmax function is used to obtain the probability distribution vector of all candidate propositions as follows:
[0031]
[0032] In the formula, [·;·] represents the vector concatenation operation;
[0033] The answer obtained by large language model reasoning is:
[0034]
[0035] Among them, the argmax function represents the parameter when the maximum value of the vector is obtained;
[0036] like Then we have:
[0037]
[0038] For any j∈[1,n], j≠i, we have:
[0039] γ i ≥γ j .
[0040] As a preferred solution, the large language model adopts the LLaMa model; the evidence theory adopts the Yager improved synthesis rule to synthesize the confidence scores of all candidate answers a, and handles the conflicts in the confidence scores of different candidate answers a by introducing a discount factor and a conflict coefficient.
[0041] As a preferred solution, the step of judging whether the answer with the highest total belief function score is consistent with the answer inferred by the large language model, and if not, modifying the confidence score of the large language model for the candidate answer a or modifying the answer inferred; iterating the loop until the answer inferred by the large language model is consistent with the answer with the highest total belief function score, comprises:
[0042] Assume that the answer with the highest total belief function score is The answer obtained by large language model reasoning is Let the large language model be based on the most supported The evidence and the most opposing If the large language model does not modify the answer, the large language model is required to regenerate the most supportive The evidence and the most opposing The score of the evidence is then retrieved, and the answer with the highest total belief function score is obtained again, and this cycle is repeated until any of the following three conditions is met:
[0043] (1) The large language model modifies the judgment so that
[0044] (2) The value of the belief function changes so that
[0045] (3) Traverse all sources of evidence;
[0046] If the loop exits with condition (3), use For the final answer;
[0047] Otherwise, let the large language model judge the problem again based on the background knowledge with the highest total belief function score, update the model's confidence score for each background knowledge based on the latest reasoning results, and re-evaluate the reasoning conclusion based on the new confidence score. Repeat the cycle until the conclusion finally inferred by the large language model is consistent with the best explanation supported by the current evidence theory (that is, the answer with the highest total belief function score).
[0048] In a second aspect, an iterative background knowledge extraction system based on a large language model is provided, comprising:
[0049] A background knowledge candidate set construction module is used to find several paragraphs in the course text that are most relevant to the questions and options as candidate knowledge sets, and to splice the candidate knowledge sets with the corresponding relationship descriptions into a background knowledge candidate set;
[0050] The confidence score acquisition module is used to convert a question q and a candidate answer a into a declarative proposition h for a sentence s in the background knowledge candidate set through a large language model, and obtain the confidence score of a sentence s and a candidate answer a of the question q;
[0051] The total belief function score screening module is used to synthesize the confidence scores of all candidate answers a using evidence theory to obtain the total belief function score of proposition h and determine the answer with the highest total belief function score;
[0052] The large language model iterative optimization output module is used to determine whether the answer with the highest total belief function score is consistent with the answer inferred by the large language model. If not, the large language model's confidence score for candidate answer a is modified or the answer inferred is modified; it iterates until the answer inferred by the large language model is consistent with the answer with the highest total belief function score, and the answer inferred by the large language model is output.
[0053] According to a third aspect, an electronic device is provided, including:
[0054] at least one processor;
[0055] and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the iterative background knowledge extraction method based on the large language model.
[0056] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the iterative background knowledge extraction method based on a large language model.
[0057] Compared with the prior art, the present invention has at least the following beneficial effects:
[0058] In order to effectively and accurately extract knowledge fragments directly related to specific questions from long textbook texts, existing textbook question-answering methods are usually difficult to handle teaching materials containing a large amount of text, especially when it is necessary to extract detailed and specific information from these texts to answer questions. In addition, existing methods are also insufficient when processing multimodal content (such as the combination of text and schematic diagrams), and cannot effectively combine these different forms of information to provide accurate answers. The present invention uses an iterative process to continuously optimize and refine the knowledge extraction process by utilizing the powerful semantic understanding ability of a large language model, thereby improving the accuracy and efficiency of the textbook question-answering system. The large language model of the present invention can accurately locate the most useful knowledge for answering questions from the long text course knowledge, and this knowledge can also serve as the basis for the final selection of a certain answer. By using a large language model, the present invention has a deeper understanding of the text content and semantics, thereby providing more accurate answers in complex textbook question-answering tasks. This deep learning method can more accurately extract information directly related to the question than traditional keyword search-based technologies. The application of large language models not only improves the semantic understanding of text, but also enables the system to handle more complex queries, including understanding long sentences, implicit meanings, and complex relationships between texts, which is particularly important for understanding and answering textbook questions that rely on deep knowledge understanding. In addition, the iterative extraction and scoring process of the present invention allows the system to continuously optimize its predictions and adjust them according to newly acquired information and context, thereby gradually improving the quality and relevance of answers in a continuous question-answering environment, and has the ability to handle large amounts of textbook content and effectively extract and analyze data from large data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0060] Figure 1 A schematic diagram of a framework of an iterative background knowledge extraction method based on a large language model according to an embodiment of the present invention;
[0061] Figure 2 A schematic diagram of an iterative background knowledge scoring update process according to an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, ordinary technicians in this field can also obtain other embodiments without making creative work.
[0063] At present, the existing large language models are still insufficient in accurately extracting knowledge fragments directly related to specific questions when processing long texts. To solve this problem, the method based on iterative background knowledge extraction continuously optimizes the extraction of knowledge through multiple rounds of iterations, and uses the intrinsic connection between questions and answers and the text scoring function of the model to accurately locate and extract key information. However, although this method improves the accuracy of information extraction, it has a high computational cost when processing large-scale text data, and has high requirements for model adjustment and optimization. In addition, these models may still have limitations in adaptability in specific fields or in understanding more complex semantics, and the efficiency and adaptability of these models need to be optimized.
[0064] The present invention designs an iterative background knowledge extraction method (IBKE) based on a large language model. The method framework is as follows: Figure 1 As shown in the figure, it mainly includes the following steps: first, extract the coarse-grained relevant content from the course text segment, then use the large language model to score the confidence of each answer for each extracted background knowledge sentence, and then use the evidence theory to synthesize these confidence scores to obtain the total belief function score of each answer. Based on the total belief function score calculated in the above steps, compare the answer with the highest score with the answer originally predicted by the large language model to determine whether the two are consistent. If they are inconsistent, the large language model will be judged again based on the background knowledge questions with high belief function scores, and the large language model will be required to modify its score for high-scoring evidence or modify the predicted answer. This process will be repeated until the predicted answer of the large language model is consistent with the answer synthesized by the evidence theory.
[0065] Specifically, the iterative background knowledge extraction method based on the large language model in the embodiment of the present invention includes the following steps:
[0066] 1. Coarse-grained candidate knowledge extraction
[0067] In the textbook question-answering data set, 75% of the course content has more than 50 sentences of course background knowledge, about one-fifth of the content has more than 100 sentences of background knowledge, and some course background knowledge even exceeds 200 sentences. If all text content is judged, while the amount of calculation is greatly increased, it may also cause evidence synthesis to be too noisy and increase uncertainty. Based on the above analysis, the embodiment of the present invention will first use the term frequency-inverse document frequency TF-IDF method to find several paragraphs in the course that are most relevant to the questions and options as candidate knowledge sets. The TF-IDF algorithm measures the importance of vocabulary by inverse document frequency. This method specifically considers the repetition frequency of stop words (StopWords) and topic vocabulary in different text paragraphs. When a word appears frequently in multiple paragraphs, the TF-IDF value will decrease, indicating that the vocabulary has a low correlation with a specific question or option. The specific content of the coarse-grained candidate knowledge extraction of the embodiment of the present invention is as follows:
[0068] For a textbook question-answering problem q∈Q, we first concatenate the question and its candidate answers, and express the concatenated result as Each item in the set represents a word, n q is the total number of words obtained after concatenation. q ={p1, p2, ..., p j}, each paragraph p in the collection j ∈C q Treated as an independent document, the score of each ti is calculated using the TF-IDF method, and then the sum is calculated to get the score of the paragraph for the question option:
[0069]
[0070] In the formula, score jq For paragraph p j For the relevance score of question q, TF-IDF is the term frequency-inverse document frequency method, specifically:
[0071] TF-IDF t,d =TF t , d ×IDF t
[0072] In the formula, TFt,d is the term frequency, which represents the number of times term t appears in document d, and IDF t is the inverse document frequency, which represents the rarity of term t in all documents, that is:
[0073]
[0074] Where N is the total number of documents, n(t) is the number of documents containing t. Specifically for the textbook question answering task, N = m q .
[0075] Based on the above formula, the three paragraphs with the highest scores are retained and concatenated with the relationship description to form a background knowledge candidate set. s q is the number of sentences in the background knowledge candidate set for question q.
[0076] Using K q You can proceed to the next step.
[0077] 2. Background knowledge scoring of large language models
[0078] The embodiment of the present invention uses the LLaMa model to output the score of background knowledge. LLaMA (Large Language Model Meta AI) is a large language model series developed by the MetaAI research team to promote research and application in the field of natural language processing (NLP). The LLaMA model has attracted attention for its small model size and efficient performance. It has shown performance that is competitive with larger models in multiple benchmarks, while being more efficient in terms of resource requirements. qIf the sentence s is regarded as a candidate source of evidence, then for a certain sentence s, the embodiment of the present invention first converts the question q and the candidate answer a into a declarative proposition h through a large language model, and then scores from three dimensions to obtain the confidence score of the sentence s and a candidate answer a of the question q. The three dimensions are: relevance, support, and sufficiency. Among them, relevance represents the degree of connection between the background knowledge and the proposition. A high degree of relevance indicates that the background knowledge of the sentence directly involves the core issue of the problem, while low-correlation background knowledge may only provide indirect or peripheral information. Assuming the proposition is: "Cactus can survive in the desert", then "Cactus has the ability to store water and a structure that reduces water evaporation, making it extremely drought-resistant." is a highly relevant sentence that provides evidence or logical support for the answer. High support means that the background knowledge can strongly support or prove the answer, while low support means that the background knowledge does not support the answer. Specifically, assuming that for the proposition "Antibiotics are ineffective against bacteria", the background knowledge "Antibiotics are a class of drugs that can kill or inhibit the growth of bacteria" has low support for the proposition, while "Some bacteria may evolve drug resistance" has a slightly higher support for the proposition. Sufficiency involves whether the background knowledge provides enough information to fully support the answer, including all necessary details and conditions. High sufficiency means that the background knowledge can independently constitute a complete support for the answer, while low sufficiency may require additional information or background knowledge to complete the answer. Take the proposition "Plants convert solar energy into chemical energy through photosynthesis." as an example. At this time, the background knowledge "Photosynthesis can consume carbon dioxide and produce oxygen" is an example of insufficient sufficiency, while "Photosynthesis can synthesize energy-rich organic matter from carbon dioxide and water" is an example of higher sufficiency.
[0079] In practice, considering that low support indicates that the model is against the proposition, and low relevance and low sufficiency only lead to increased uncertainty, the score range of support is set to [-1, 1], while relevance and sufficiency are set to (0, 1). The final score is recorded as E sqa , support, relevance and sufficiency scores are denoted as e s , e r , e a ,have:
[0080] E sqa =max(e s ×e r ×e a ,0)
[0081] In the formula, E sqa is the belief function value of sentence s for the proposition h composed of question q and answer a. For convenience, we agree that E sqa =E sh, where proposition h consists of question q and answer a.
[0082] 3. Evidence Theory Answer Prediction
[0083] All scores of all candidate evidence are combined using evidence theory to obtain the final score. Evidence theory is a mathematical framework for dealing with uncertainty and the combination of evidence. It was independently proposed by Glenn Shafer and Arthur Dempster to address the limitations of probability theory in dealing with uncertainty. The main steps of evidence theory are as follows:
[0084] First, define the propositions and the evidence related to each proposition, then obtain the belief function value of each evidence for each proposition, then use the combination rule to combine the scores of all evidence for each proposition, and finally predict the answer based on the total score. In this embodiment of the present invention, the proposition h is defined as the declarative sentence transformed from the question q and the candidate answer a, and the evidence source set is the background knowledge candidate set K q ; At the same time, the belief function value is also determined by E sqa Calculated, therefore, the final result can be predicted by combining the evidence scores of each proposition using the combination rule. Since the traditional Dempster-Shafer synthesis rule has some problems in calculation, namely, it is difficult to handle conflicting evidence and is sensitive to basic probability assignments, the embodiment of the present invention adopts the Yager improved synthesis rule to combine the evidence scores of each proposition. This rule is an optimization of the traditional Dempster-Shafer synthesis rule, and flexibly handles conflicts between different evidence scores by introducing discount factors and conflict coefficients. The specific algorithm is as follows:
[0085] Table 1
[0086]
[0087] The algorithm of the embodiment of the present invention is based on the evidence source set K q The belief function score dictionary M for the proposition set H v As input, specifically, for a specific source of evidence k and a proposition h consisting of q and a, the following relationship exists:
[0088] M v [k][h]=E sqa
[0089] The output of the algorithm is the total belief function score Mc[h] for each proposition h∈H. On this basis, after splicing the total belief function scores of all propositions, the softmax function can be used to obtain the probability distribution vector of all candidate propositions.
[0090]
[0091]
[0092] In the formula, [·;·] represents the vector concatenation operation. At this time, the correct answer predicted by the model is:
[0093]
[0094] in, is the final prediction result. The argmax function represents the parameter when the maximum value of the vector is obtained. That is, if have:
[0095]
[0096] For any j∈[1,n], j≠i, we have:
[0097] γ i ≥γ j
[0098] 4. Iteratively update background knowledge scoring
[0099] In order to avoid as much as possible the problems existing in the large language model itself that may cause the generated score to be inaccurate, the embodiment of the present invention proposes a method for iteratively updating the background knowledge score by using the large language model. The specific algorithm is as follows:
[0100] Table 2
[0101]
[0102] Specifically, after obtaining the total belief function score of each proposition, let the large language model directly judge the proposition based on all evidence sources, and check whether the answer with the highest total belief function score is consistent with the answer judged by the large language model. If it is inconsistent, it means there are two possible situations. One is that the judgment of the large language model is wrong, and the other is that one or several scorings of the large language model are wrong. Suppose that the answer with the highest total belief function score is The answer directly determined by the large language model is At this time, let the large language model be based on the most supported The evidence and the most opposing If the large language model does not modify the answer, the large language model is required to regenerate the most supportive The evidence and the most opposing The score of the evidence is then retrieved to find the answer with the highest total belief function score.
[0103] This cycle continues until any of the following three conditions is met:
[0104] (1) The large language model modifies the judgment so that
[0105] (2) The value of the belief function changes so that
[0106] (3) Traverse all sources of evidence;
[0107] If the loop exits with condition (3), use For the final answer.
[0108] Otherwise, let the large language model judge the question again based on the background knowledge with high belief function values, adjust the confidence score of the large language model on the relevant background knowledge according to the belief function score of the high-scoring evidence, or re-evaluate the model's predicted answer, and require the large language model to modify its score on the high-scoring evidence or modify the predicted answer. This process will be repeated until the predicted answer of the large language model is consistent with the answer with the highest total belief function score synthesized by the evidence theory, so as to more accurately reflect the support level of the current evidence. In summary, the iterative background knowledge extraction method based on the large language model can be summarized into the following stages: First, the TF-IDF method is used to perform coarse-grained extraction of background courses, and spliced with possible schematic diagram relationship descriptions to obtain a candidate evidence source set K q ; Then use the large language model to calculate each proposition h and evidence s∈K q The belief function score E sh , and store it in the dictionary M v [k][h]=E sh ; Then use evidence theory to calculate the total belief function score of each proposition, and finally use step 4 to iteratively update the large language model.
[0109] The effect of the iterative background knowledge extraction method based on the large language model in the embodiment of the present invention is analyzed through comparative experiments.
[0110] The embodiment of the present invention conducts relevant experiments on the textbook question-answering dataset. In the schematic diagram problem, the iterative background knowledge extraction method (IBKE) based on the large language model is integrated with the structure parsing network (SPN), and compared with other models on this basis. The evaluation index is accuracy, and the experimental results are shown in Table 3.
[0111] Table 3
[0112]
[0113] A detailed comparison of the impact of several configurations, including the LLaMa-13B model, on the textbook question answering task, and how the performance of these models compares to existing technologies. The results show that the LLaMa-13B model exhibits high accuracy and relevance. The model uses a large language model and an iterative extraction method to improve the performance of textbook question answering by accurately locating and extracting knowledge fragments that are directly related to the question. The comparison results show that LLaMa-13B outperforms LLaMa-7B and other baseline methods in multiple evaluation indicators, such as answer generation accuracy, relevance scoring, and question answering efficiency.
[0114] The comparative experiment also involved testing different extraction strategies, including using IR (information retrieval) and TF-IDF technology as comparison baselines. The results show that the LLaMa model can more effectively utilize the capabilities of the language model when integrating these technologies for background knowledge extraction. Compared with the traditional method alone, it has obvious advantages in understanding and generating answers.
[0115] The analysis shows that (1) for the zero-shot baseline model that also uses LLaMa-13B, the structural parsing network based on the synthetic schematic dataset and the iterative background knowledge extraction method based on the large language model (SPN+IBKE) proposed in the embodiment of the present invention are significantly better than the traditional information parsing method, and achieve approximately 2.5%, 3.8% and 2.4% improvement in the judgment questions, multiple-choice questions and schematic choice questions respectively; (2) for the baseline model fine-tuned on the textbook question-answering dataset, the method proposed in the embodiment of the present invention surpasses most supervised methods, but is still some distance away from the best method currently, which may be because it has been fine-tuned multiple times on several additional datasets (such as RACE, ARC, OpenbookQA, etc.).
[0116] Therefore, compared with the existing textbook methods, the embodiments of the present invention have the following advantages: (1) Compared with the zero-sample baseline, the method of the embodiment of the present invention adopts an iterative update strategy to repeatedly update the output of the large language model to maximize the avoidance of answer judgment problems caused by the randomness of the output and the uncertainty of the information; (2) Compared with the model fine-tuned on the data set, the method of the present invention does not require training, has no dependence on training data, and can be applied to all textbook question-answering tasks. At the same time, it also has the following shortcomings: Although the model effect of the embodiment of the present invention surpasses all zero-sample baseline models and most supervised baseline models, it still has a certain effect gap compared with the models that have been additionally fine-tuned on other data sets; at the same time, the module of the present invention is highly dependent on the performance of the large language model. If the performance of the large language model itself is insufficient, the results will be reduced. As shown in Table 4, the results of changing the language model from LLaMa-13B to LLaMa-7B show that changing the language model brings a 2.0% to 3.2% effect loss. Therefore, the following future improvement directions can be proposed:
[0117] First, we can consider changing the scoring method and try to design a differentiable confidence score to make the process optimizable, which means extending the work to the supervised training setting. Second, we can consider replacing the large language model with a better effect or using domain data to fine-tune the large model to increase the model's ability to understand the text content and thus improve the model's performance.
[0118] Table 4
[0119]
[0120] Ablation experiment analysis
[0121] Table 5 shows the ablation experiment results, which are used to verify the effectiveness of the model and each module of the embodiment of the present invention. Considering that some modules are only for schematic diagrams, the ablation experiment is conducted on schematic diagram selection questions.
[0122] Table 5
[0123]
[0124] In Table 5, ALL represents the initial model without modification, -SPN represents the abandonment of the structural analysis relationship generation of the schematic diagram in the research content 1, -Diagram represents the abandonment of all schematic diagram related content, -IUKE represents the direct use of the synthesis result without iterative update, and -E sqaIt means that the large language model directly generates the belief function score, and -ALL means that the LLaMa model is used directly without any means to change the input. As can be seen from Table 5, after deleting the SPN module, the model results decreased, indicating that the relationship description generated by the SPN module has a certain improvement on the present invention; after deleting all schematic contents, the accuracy rate decreased significantly, indicating that when answering schematic questions, schematic information is more important; the decrease in the result of deleting the multi-dimensional score proves the effectiveness of the scoring strategy; the large decrease in the result after deleting the iterative update shows that the output of the large language model is indeed uncertain; the experimental results using only the LLaMa model show that the model cannot solve the textbook question-answering task without processing the relevant knowledge.
[0125] These experimental results emphasize the effectiveness and superiority of the iterative background knowledge extraction method based on the large language model in modern textbook question-answering applications, especially in dealing with problems involving deep knowledge understanding and long text processing. Through these experiments, it can be seen that the embodiment of the present invention can use the large language model to score course-related texts to solve the problem of insufficient annotation information, and on the other hand, it uses an iterative method to make the large language model check and update its own output.
[0126] Another embodiment of the present invention further provides an iterative background knowledge extraction system based on a large language model, comprising:
[0127] A background knowledge candidate set construction module is used to find several paragraphs in the course text that are most relevant to the questions and options as candidate knowledge sets, and to splice the candidate knowledge sets with the corresponding relationship descriptions into a background knowledge candidate set;
[0128] The confidence score acquisition module is used to convert a question q and a candidate answer a into a declarative proposition h for a sentence s in the background knowledge candidate set through a large language model, and obtain the confidence score of a sentence s and a candidate answer a of the question q;
[0129] The total belief function score screening module is used to synthesize the confidence scores of all candidate answers a using evidence theory to obtain the total belief function score of proposition h and determine the answer with the highest total belief function score;
[0130] The large language model iterative optimization output module is used to determine whether the answer with the highest total belief function score is consistent with the answer inferred by the large language model. If not, the large language model's confidence score for candidate answer a is modified or the answer inferred is modified; it iterates until the answer inferred by the large language model is consistent with the answer with the highest total belief function score, and the answer inferred by the large language model is output.
[0131] Another embodiment of the present invention further provides an electronic device, including:
[0132] at least one processor;
[0133] and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the iterative background knowledge extraction method based on the large language model.
[0134] Another embodiment of the present invention further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the iterative background knowledge extraction method based on a large language model.
[0135] Exemplarily, the instructions stored in the memory may be divided into one or more modules / units, which are stored in a computer-readable storage medium and executed by the processor to complete the iterative background knowledge extraction method based on a large language model of the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the server.
[0136] The electronic device may be a computing device such as a smart phone, a notebook, a PDA, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that the electronic device may also include more or fewer components, or a combination of certain components, or different components, for example, the electronic device may also include an input / output device, a network access device, a bus, etc.
[0137] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0138] The memory may be an internal storage unit of the server, such as a hard disk or memory of the server. The memory may also be an external storage device of the server, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the server. Furthermore, the memory may include both an internal storage unit of the server and an external storage device. The memory is used to store the computer-readable instructions and other programs and data required by the server. The memory may also be used to temporarily store data that has been output or is to be output.
[0139] It should be noted that the information interaction, execution process and other contents between the above-mentioned module units are based on the same concept as the method embodiment. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0140] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0141] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk.
[0142] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0143] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An iterative background knowledge extraction method based on a large language model, characterized in that: include: Find several paragraphs in the course text that are most relevant to the questions and options as candidate knowledge sets, and combine the candidate knowledge sets with the corresponding relationship descriptions to form a background knowledge candidate set; For a sentence s in the background knowledge candidate set, the question q and the candidate answer a are converted into a declarative proposition h through the large language model, and the confidence score of the sentence s and the candidate answer a of the question q is obtained; Use evidence theory to synthesize the confidence scores of all candidate answers a to obtain the total belief function score of proposition h, and determine the answer with the highest total belief function score; Determine whether the answer with the highest total belief function score is consistent with the answer inferred by the large language model. If not, modify the confidence score of the large language model for candidate answer a or modify the answer inferred. Iterate until the answer inferred by the large language model is consistent with the answer with the highest total belief function score, and output the answer inferred by the large language model.
2. The iterative background knowledge extraction method based on a large language model according to claim 1, characterized in that: The step of finding several paragraphs most relevant to the questions and options in the course text as candidate knowledge sets is implemented using the term frequency-inverse document frequency TF-IDF method, which measures the importance of vocabulary by using the inverse document frequency, combined with the stop words StopWords and the repetition frequency of topic vocabulary in different text paragraphs. When a word appears multiple times in multiple paragraphs, the TF-IDF value decreases, indicating that the corresponding word has a low correlation with the specific question or option.
3. The iterative background knowledge extraction method based on a large language model according to claim 2, characterized in that: The step of searching for several paragraphs in the course text that are most relevant to the questions and options as candidate knowledge sets, and splicing the candidate knowledge sets with the corresponding relationship descriptions into a background knowledge candidate set includes: For a textbook question-answering problem q∈Q, concatenate the question and the corresponding answer candidates, and express the concatenated result as Each item in the set represents a word, n q is the total number of words obtained after concatenation; For the paragraph set C in the problem background course q ={p1, p2, ..., p j }, each paragraph p in the collection j ∈C q It is regarded as an independent document. The score of each word ti is calculated by the term frequency-inverse document frequency TF-IDF method. The sum is used to get the score of the corresponding paragraph for the question option. The calculation expression is as follows: In the formula, Score jq For paragraph p j Relevance score for question q; TF-IDF t,d =TF t,d ×IDF r Where TFt,d is the term frequency, which represents the number of times term t appears in document d; IDF t is the inverse document frequency, which represents the rarity of term t in all documents. The calculation expression is as follows: Where N is the total number of documents, n (t) is the number of documents containing t; Sort by score, retain a number of paragraphs, and concatenate them with the relationship description to form a background knowledge candidate set s q is the number of sentences in the background knowledge candidate set for question q.
4. The iterative background knowledge extraction method based on a large language model according to claim 3, characterized in that: In the step of obtaining the confidence score of a sentence s and a candidate answer a to question q, the confidence score is performed from three dimensions: relevance, support, and sufficiency; wherein relevance represents the degree of connection between the background knowledge and the proposition; support represents whether the background knowledge can support the corresponding answer; sufficiency represents whether the background knowledge provides enough information to fully support the answer, including all necessary details and conditions; The support score range is set to [-1, 1], the relevance and sufficiency score range is set to (0, 1]; the confidence score is recorded as E sqa , the support, relevance and sufficiency scores are denoted as e s , e r , e a ,but: AND sqa =max(e s ×and r ×and a ,0) In the formula, the confidence score E sqa is the belief function value of sentence s for the proposition h consisting of question q and candidate answer a.
5. The iterative background knowledge extraction method based on a large language model according to claim 4, characterized in that: The confidence scores of all candidate answers a are synthesized using evidence theory to obtain the total belief function score of proposition h, and the answer with the highest total belief function score is determined, including: Evidence theory defines the background knowledge candidate set K q As a set of candidate evidence sources, the belief function score dictionary M of the candidate evidence source set on the proposition set H v As input, for a particular source of evidence k and a proposition h consisting of q and a, there exists: M v [k][h]=E sqa The output of the evidence theory is the total belief function score Mc[j] for each proposition h∈H; after concatenating the total belief function scores of all propositions, the softmax function is used to obtain the probability distribution vector of all candidate propositions as follows: In the formula, [·;·] represents the vector concatenation operation; The answer obtained by large language model reasoning is: Among them, the argmax function represents the parameter when the maximum value of the vector is obtained; like Then we have: For any j∈[1,m], j≠i, we have: c i ≥c j 。 6. The iterative background knowledge extraction method based on a large language model according to claim 1, characterized in that: The large language model adopts the LLaMa model; the evidence theory adopts the Yager improved synthesis rule to synthesize the confidence scores of all candidate answers a, and handles the conflict of confidence scores of different candidate answers a by introducing a discount factor and a conflict coefficient.
7. The iterative background knowledge extraction method based on a large language model according to claim 1, characterized in that: The step of determining whether the answer with the highest total belief function score is consistent with the answer inferred by the large language model, and if not, modifying the confidence score of the large language model for the candidate answer a or modifying the answer inferred; The steps of iterating the loop until the answer obtained by the large language model inference is consistent with the answer with the highest score of the total belief function include: Assume that the answer with the highest total belief function score is The answer obtained by large language model reasoning is Let the large language model be based on the most supported The evidence and the most opposing If the large language model does not modify the answer, the large language model is required to regenerate the most supportive The evidence and the most opposing The score of the evidence is then retrieved, and the answer with the highest total belief function score is obtained again, and this cycle is repeated until any of the following three conditions is met: (1) The large language model modifies the judgment so that (2) The value of the belief function changes so that (3) Traverse all sources of evidence; If the loop exits with condition (3), use For the final answer; Otherwise, based on the background knowledge with the highest total belief function score, the large language model is asked to re-evaluate the confidence score of the relevant evidence or re-infer the answer, and this process is iterated until the predicted answer of the large language model is consistent with the answer with the highest total belief function score synthesized by the evidence theory.
8. An iterative background knowledge extraction system based on a large language model, characterized in that: include: A background knowledge candidate set construction module is used to find several paragraphs in the course text that are most relevant to the questions and options as candidate knowledge sets, and to splice the candidate knowledge sets with the corresponding relationship descriptions into a background knowledge candidate set; The confidence score acquisition module is used to convert a question q and a candidate answer a into a declarative proposition h for a sentence s in the background knowledge candidate set through a large language model, and obtain the confidence score of a sentence s and a candidate answer a of the question q; The total belief function score screening module is used to synthesize the confidence scores of all candidate answers a using evidence theory to obtain the total belief function score of proposition h and determine the answer with the highest total belief function score; The large language model iterative optimization output module is used to determine whether the answer with the highest total belief function score is consistent with the answer inferred by the large language model. If not, the large language model's confidence score for candidate answer a is modified or the answer inferred is modified; it iterates until the answer inferred by the large language model is consistent with the answer with the highest total belief function score, and the answer inferred by the large language model is output.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the iterative background knowledge extraction method based on a large language model as described in any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the iterative background knowledge extraction method based on a large language model as described in any one of claims 1-7.
Citation Information
Cited By
Scientific and technical literature question and answer method and system based on self-feedback iterative optimization
CN120950660A