Activity assessment method for factual knowledge question, product, equipment and medium

By collecting historical knowledge graphs in large language models, generating a set of factual knowledge problems, and conducting prior judgments and posterior inspections, the problem of insufficient metacognitive ability in the big model in determining whether it knows certain facts is solved, improving the efficiency of answering ability evaluation and knowledge cognitive ability, and preventing misjudgments caused by hallucinations.

CN120180060AActive Publication Date: 2025-06-20LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Patent Information

Application Number
CN202510668928.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Large language models have insufficient metacognitive ability in determining whether they know certain facts, which leads to the possibility of giving wrong answers, which is called the big model hallucination.

Method used

By collecting historical knowledge graphs from the target field, a factual knowledge problem evaluation set is generated, and a priori judgment and posterior inspection are performed using a pretrained language model. A priori judgment reasons on factual knowledge questions through the preset priori prompt information group. If the inference answer does not contain abnormal characters, a posterior check is performed to verify the authenticity of the answer.

Benefits of technology

The efficiency of the pre-trained language model's answering ability to factual knowledge questions is improved, the model's cognitive ability of its own knowledge state is enhanced, and misjudgment caused by hallucinations is effectively prevented.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180060A_ABST
    Figure CN120180060A_ABST
Patent Text Reader

Abstract

The invention discloses a factual knowledge question answering ability evaluation method, product, equipment and medium, and relates to the field of natural language processing, and the method comprises the steps: sampling sub-graphs in a historical knowledge graph to generate a factual knowledge question evaluation set; inputting the evaluation set and the prior prompt information group into a pre-training language model so as to perform reasoning on the factual knowledge problem to obtain a plurality of reasoning answers; and if the plurality of reasoning answers do not contain the preset target character, detecting the authenticity of the plurality of reasoning answers by using posterior prompt information, and if the detection result shows that the plurality of reasoning answers are consistent with the fact, judging that the model has the capability of answering the corresponding question. According to the method, the factual knowledge problem is reasoned at first through two stages of prior judgment and posterior check, and then authenticity detection is performed on the reasoned answer, so that misjudgment of the model caused by illusion can be prevented, and the evaluation efficiency of the model on the answering ability of the factual knowledge problem is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing, and particularly to methods, products, devices, and media for evaluating the ability to answer factual knowledge questions. Background Art

[0002] Currently, the rapid development of large language models (LLMs) has brought revolutionary progress to the fields of natural language processing (NLP) and artificial intelligence, but also faces some challenges and problems. Through the training of massive texts, large language models (which can be applied to multiple fields such as education, healthcare, finance, law, and content creation) have mastered a large amount of factual knowledge. However, large language models are relatively lacking in the ability to judge whether they know certain facts. For example, when a human faces a question, they can judge whether they have the ability to answer the question, but for large models, this metacognitive ability is relatively weak. Therefore, when a user asks a question, the large language model may give an answer that seems reasonable but is actually incorrect, that is, the large model hallucination.

[0003] To solve the above problems, the following two methods have been proposed currently: One method is to introduce external knowledge to help the large language model evaluate its own ability. For example, a knowledge base is provided to the model, and the model can query the knowledge base to judge whether it has the ability to answer the user's question. However, this method depends on the knowledge base, and if the corresponding content cannot be found in the knowledge base, a correct judgment cannot be made. Another method is to use learning from examples, providing the model with some examples with known answers, and letting the model learn to judge whether it can correctly answer similar questions. However, this method does not explicitly introduce metacognitive ability. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide methods, products, devices, and media for evaluating the ability to answer factual knowledge questions, which can improve the efficiency of evaluating the ability of pre-trained language models to answer factual knowledge questions and the model's cognitive ability of its own knowledge state. The specific solutions are as follows: In the first aspect, this application discloses a method for evaluating the ability to answer factual knowledge questions, including: Collecting the historical knowledge graph of the target field and sampling the subgraphs in the historical knowledge graph to generate a factual knowledge question evaluation set for the target field; Input the factual knowledge question evaluation set and the preset prior prompt information group into the pre-trained language model, and infer the factual knowledge questions in the factual knowledge question evaluation set one by one based on multiple prompt statements in the prior prompt information group through linear verification to obtain multiple inference answers corresponding to the factual knowledge questions; If none of the multiple inference answers contain the preset target character, input the multiple inference answers and the preset posterior prompt information into the pre-trained language model to detect the authenticity of the multiple inference answers using the posterior prompt information, and obtain the detection result; the preset target character is a character indicating that the inference answer is abnormal; If the detection result shows that the multiple inference answers are consistent with the facts, it is determined that the pre-trained language model has the ability to answer the corresponding factual knowledge questions in the factual knowledge question evaluation set.

[0005] In a second aspect, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned factual knowledge question answering ability evaluation methods are implemented.

[0006] In a third aspect, the present application discloses an electronic device, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the above-mentioned factual knowledge question answering ability evaluation method is implemented.

[0007] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the above-mentioned factual knowledge question answering ability evaluation method is implemented.

[0008] It can be seen that in this application, the historical knowledge graph of the target domain is first collected, and sub-graphs in the historical knowledge graph are sampled to generate a factual knowledge question evaluation set for the target domain. Then, the factual knowledge question evaluation set and a preset group of prior prompt information are input into a pre-trained language model to sequentially reason about the factual knowledge questions in the factual knowledge question evaluation set based on multiple prompt statements in the group of prior prompt information through linear verification, and multiple reasoning answers corresponding to the factual knowledge questions are obtained. If none of the multiple reasoning answers contain a preset target character indicating that the reasoning answer is abnormal, the multiple reasoning answers and the preset posterior prompt information are input into the pre-trained language model to detect the authenticity of the multiple reasoning answers using the posterior prompt information to obtain a detection result. If the detection result indicates that the multiple reasoning answers are consistent with the facts, it is determined that the pre-trained language model has the ability to answer the corresponding factual knowledge questions in the factual knowledge question evaluation set. When evaluating the ability of the pre-trained language model to answer factual knowledge questions, this application is divided into two stages, specifically including two stages: prior judgment and posterior inspection. The first stage is to create a factual knowledge question evaluation set based on the historical knowledge graph of a certain domain, and then sequentially reason about each factual knowledge question in the factual knowledge question evaluation set using multiple prompt statements in the preset group of prior prompt information. The second stage is to input the multiple reasoning answers corresponding to each factual knowledge question obtained by reasoning and the preset posterior prompt information into the pre-trained language model to detect the authenticity of the multiple reasoning answers. If the detection result indicates that the multiple reasoning answers are consistent with the facts, it is determined that the model has the ability to answer the corresponding factual knowledge questions. Through the above two stages (i.e., the prior judgment stage and the posterior inspection stage), as well as the prompt information preset in each stage, not only can the model be guided to judge whether it has the ability to answer the corresponding factual knowledge questions, but also through the posterior inspection operation, the model can be effectively prevented from misjudging factual knowledge questions due to hallucinations. In addition, through the above method, the detection of the model's factual knowledge answering ability can be realized without introducing additional model training computational complexity, thereby improving the evaluation efficiency of the ability to answer factual knowledge questions and the model's cognitive ability of its own knowledge state. Description of the Drawings

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0010] Figure 1Flowchart of a method for evaluating the answering ability of factual knowledge questions disclosed in this application; Figure 2 Block diagram of a process for evaluating the answering ability of a specific factual knowledge question disclosed in this application; Figure 3 Flowchart of a method for evaluating the answering ability of a specific factual knowledge question disclosed in this application; Figure 4 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners

[0011] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0012] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0013] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0014] The embodiments of the present application disclose a method for evaluating the answering ability of factual knowledge questions. Refer to Figure 1 As shown, the method includes: Step S11: Collect the historical knowledge graph of the target domain, and sample the sub-graphs in the historical knowledge graph to generate a factual knowledge question evaluation set for the target domain.

[0015] In this embodiment, first, the knowledge graphs (such as Wikidata, Freebase, etc.) in the fields that need to be evaluated for answering ability (such as fields of education, healthcare, finance, law, content creation, etc.) are collected to obtain the historical knowledge graphs of the corresponding fields. Then, subgraphs in the historical knowledge graphs of the fields are sampled to generate a fact-based knowledge question evaluation set for the fields. As can be seen from the above, the fact-based knowledge question evaluation set in this application is created in real time and dynamically based on the historical knowledge graphs. Compared with directly using existing fact-based knowledge question evaluation sets, more accurate and targeted questions can be asked for the corresponding fields. Among them, the knowledge graph (KG) in any field needs to include entities, relationships, and other contents.

[0016] Specifically, sampling the subgraphs in the historical knowledge graph to generate a fact-based knowledge question evaluation set for the target field may include: sampling the subgraphs in the historical knowledge graph based on different sampling rules to obtain multiple subgraph sets; merging the multiple subgraph sets to obtain an evaluation subgraph set, and inputting the evaluation subgraph set into a fact-based knowledge question generation template created in advance to construct a fact-based knowledge question evaluation set covering the evaluation subgraph set. That is, when sampling the subgraphs in the historical knowledge graph, sampling can be performed based on different pre-set sampling rules to obtain subgraph sets under each sampling rule, and then the subgraph sets under all sampling rules are merged, and the merged evaluation subgraph set is input into the fact-based knowledge question generation template created in advance, so as to construct a fact-based knowledge question evaluation set covering the entire evaluation subgraph set.

[0017] In a specific implementation manner, sampling the subgraphs in the historical knowledge graph based on different sampling rules to obtain multiple subgraph sets may specifically include: sampling the subgraphs in the historical knowledge graph based on a high-frequency knowledge sampling strategy and a long-tail knowledge sampling strategy to obtain a high-frequency subgraph set and a long-tail subgraph set; where the high-frequency knowledge sampling strategy is a sampling strategy that takes the nodes with the top pre-set number of node degrees among the nodes in the knowledge graph as high-frequency knowledge nodes; the long-tail knowledge sampling strategy is a sampling strategy that takes the nodes in the subgraphs of the knowledge graph with subgraph entropy values greater than the pre-set entropy value as long-tail knowledge nodes. In this embodiment, in order to ensure coverage of the knowledge graph, both high-frequency knowledge and long-tail knowledge are considered when sampling the subgraphs in the historical knowledge graph. Specifically, the high-frequency knowledge sampling strategy can be first used to sample the subgraphs in the historical knowledge graph to obtain the corresponding high-frequency subgraph set , where the high-frequency knowledge sampling strategy specifically refers to the sampling strategy of taking the nodes with the top preset number of node degrees (such as the top K nodes, with K = 1000 by default) among the nodes of each node in the knowledge graph as high-frequency knowledge nodes. That is, the high-frequency knowledge sampling strategy is based on the degree of knowledge nodes. The higher the degree, the stronger the relevance of the node in the knowledge graph, representing high-frequency knowledge. Among them, the calculation formula of the node degree is: ; In the formula, is the degree of node v, representing the total number of edges (i.e., relationships) associated with the node in the knowledge graph; is the set of edges in the knowledge graph, and each edge represents the relationship between two nodes (such as "Li Bai - Dynasty - Tang Dynasty"); is the indicator function, which takes the value of 1 when node v belongs to edge e, and 0 otherwise.

[0018] Next, sample the subgraphs in the historical knowledge graph based on the long-tail knowledge sampling strategy to obtain the corresponding set of long-tail subgraphs ; among them, the long-tail knowledge sampling strategy is to take the nodes in the subgraphs of the knowledge graph with subgraph entropy values greater than the preset entropy value as the sampling strategy for long-tail knowledge nodes. That is, the long-tail knowledge sampling strategy is based on graph entropy. The purpose of this strategy is to capture low-frequency but key diverse knowledge. Subgraphs with high entropy values contain more low-frequency but diverse nodes, representing long-tail knowledge. Among them, the calculation formula of the subgraph entropy value is: ; In the formula, represents the entropy value of subgraph S, which is used to measure the diversity of node distribution. The larger the entropy value, the more dispersed the node distribution; represents the occurrence probability of node v in subgraph S, and the calculation formula is: .

[0019] The long-tail knowledge sampling strategy is to select low-degree nodes with subgraph entropy values , where the preset entropy value can be obtained based on experimental verification and can be adjusted according to actual application requirements for screening long-tail knowledge nodes, and can be defaulted to 2.5.

[0020] It should be noted that the high-frequency knowledge sampling strategy is used to screen the core nodes with high relevance in the knowledge graph (such as "Scientist M"), covering common sense knowledge. The long-tail knowledge sampling strategy is used to capture low-frequency but key knowledge nodes (such as "quantum entanglement experimental device"), which can improve the comprehensiveness of the evaluation. By sampling the subgraphs in the knowledge graph of a certain field through a hybrid sampling strategy based on the high-frequency knowledge sampling strategy and the long-tail knowledge sampling strategy, the coverage of the knowledge graph can be ensured, and a high-frequency subgraph set and a long-tail subgraph set can be obtained. Furthermore, a factual knowledge question evaluation set covering high-frequency knowledge and long-tail knowledge can be generated, solving the problem of coverage deviation in the static evaluation set, which is beneficial to the accuracy of the subsequent evaluation of the pre-trained language model's ability to answer factual knowledge questions.

[0021] Specifically, multiple subgraph sets are merged to obtain an evaluation subgraph set, and the evaluation subgraph set is input into a pre-created factual knowledge question generation template to construct a factual knowledge question evaluation set covering the evaluation subgraph set, which may include: merging the high-frequency subgraph set and the long-tail subgraph set to obtain an evaluation subgraph set, and inputting the evaluation subgraph set into a pre-created factual knowledge question generation template to construct a factual knowledge question evaluation set covering the evaluation subgraph set. In this embodiment, first, the high-frequency subgraph set corresponding to the target domain and the long-tail subgraph set are merged to obtain an evaluation subgraph set , and then the evaluation subgraph set is input into a pre-created factual knowledge question generation template, thereby constructing a factual knowledge question evaluation set that can cover the entire evaluation subgraph set

[0022] It should be noted that the types of factual knowledge question generation templates specifically include direct question type, triple masking type, and path query type; among them, the direct question type is the type that directly generates questions based on triples in the knowledge graph; the triple masking type is the type that generates fill-in-the-blank questions by randomly masking the head entity or tail entity; the path query type is the type that extracts relationship paths with a length less than or equal to a preset length and generates multi-hop reasoning questions based on the relationship paths. Specifically, for the direct question type: for the triple (h, r, t), generate the question "What is the r that h belongs to?" For example, when the triple (h, r, t) is (Li Bai, dynasty, Tang Dynasty), the generated question is "What is the dynasty that Li Bai belongs to?"; for the triple masking type: by randomly masking the head entity h or the tail entity t, generate a fill-in-the-blank question. For example, when the triple (h, r, t) is (Li Bai, dynasty, Tang Dynasty), generate "____'s dynasty is the Tang Dynasty"; for the path query type: extract relationship paths with a length ≤ 3 and generate multi-hop reasoning questions based on the relationship paths, such as "Li Bai → 'Looking at Lushan Waterfall' → What is the dynasty of the proposer?". By using multiple pre-set factual knowledge question generation templates, different types of factual knowledge questions can be generated, which is beneficial for the model to comprehensively evaluate factual knowledge questions, thereby improving the evaluation accuracy of the model's answering ability for factual knowledge questions.

[0023] In this embodiment, the evaluation sub-graph set is input into the pre-created factual knowledge question generation template to construct a factual knowledge question evaluation set covering the evaluation sub-graph set, which specifically includes: inputting the evaluation sub-graph set into the pre-created factual knowledge question generation template including direct question type, triple masking type, and path query type to generate the corresponding number of questions for the direct question type, triple masking type, and path query type according to a preset generation ratio, so as to obtain a factual knowledge question evaluation set covering the evaluation sub-graph set; among them, the number of questions generated by each sub-graph in the evaluation sub-graph set is determined based on the number of nodes of the corresponding sub-graph, and the ratio of the number of questions generated by the high-frequency sub-graph set to the number of questions generated by the long-tail sub-graph set is a preset value. In this embodiment, when generating a factual knowledge question evaluation set covering the evaluation sub-graph set based on the factual knowledge question generation template it is possible to generate factual knowledge questions corresponding to different template types according to a preset ratio. For example, generate the corresponding number of factual knowledge questions according to the ratio of 2:4:4 for the direct question type, triple masking type, and path query type. In addition, the number of questions generated by each sub-graph in the evaluation sub-graph set is determined based on the number of nodes of the corresponding sub-graph, and the high-frequency sub-graph set the number of questions generated by the long-tail sub-graph set When this is the case, the number of questions generated by the sub-graph is , and the distribution of the generated questions is as follows: the set of high-frequency sub-graphs generates 70% of the total generated questions, and the set of long-tail sub-graphs generates 30% of the total generated questions. That is to say, the ratio of the questions generated by the set of high-frequency sub-graphs to the questions generated by the set of long-tail sub-graphs is 7:3. By combining the structural characteristics of the knowledge graph, this application pre-designs different types of factual knowledge question generation templates and automatically constructs factual knowledge questions that can well cover the sub-graph set according to the preset proportional relationship, so as to generate a more reasonable and comprehensive factual knowledge question evaluation set for a certain field, which is beneficial to the subsequent model to more comprehensively evaluate the factual knowledge questions in the corresponding field, thereby improving the accuracy of the model's evaluation of the answering ability of the factual knowledge questions in this field.

[0024] Step S12: Input the factual knowledge question evaluation set and the preset prior hint information group into the pre-trained language model, and sequentially reason about the factual knowledge questions in the factual knowledge question evaluation set based on multiple hint statements in the prior hint information group in a linear verification manner to obtain multiple reasoning answers corresponding to the factual knowledge questions.

[0025] In this embodiment, after sampling the sub-graphs in the historical knowledge graph to obtain a factual knowledge question evaluation set for the target field, further, the factual knowledge questions in the factual knowledge question evaluation set and the preset prior hint information group can be input into the pre-trained language model (Pretrained Language Model, PLM) together, so as to perform reasoning on each factual knowledge question in the input factual knowledge question evaluation set in a linear verification manner and based on multiple hint statements in the prior hint information group, thereby obtaining multiple reasoning answers corresponding to a single factual knowledge question; where the pre-trained language model is a large language model (LLM, Large Language Model) related to the target field corresponding to the factual knowledge question evaluation set, such as a medical Q&A large model, a financial large model, a government affairs large model, etc., and this model can adopt a Transformer network structure.

[0026] In addition, the number of multiple reasoning answers corresponding to a single factual knowledge question is the same as the number of multiple hint statements in the prior hint information group. This application does not specifically limit the number of hint statements in the prior hint information group, which can be set according to actual application requirements.

[0027] For example, various factual knowledge questions and a preset group of prior prompt information (prompt group) are input into the medical Q&A large model to determine whether it knows a factual knowledge question related to a certain medical field through the medical Q&A large model. Among them, the group of prior prompt information (prompt group) specifically includes the following 4 prior prompt statements: 1. Do you know the answer to the following question? If not, please reply "I don't know". 2. Are you capable of answering the following question? If not, please reply "I don't know". 3. Do you have any knowledge of the following question? If not, please answer "I don't know". 4. Do you think you have enough wisdom to answer the following question? If not, please answer "I don't know". It should be noted that during the reasoning process of the model, a linear verification method is adopted (this method simulates the stress test scenario in human conversations), that is, for question q, the above 4 prior prompt statements are used for reasoning one by one. For example, when the reasoning result corresponding to the first prior prompt statement does not contain the negative answer character "I don't know", that is, the reasoning answer corresponding to the first prior prompt statement is normally answered, then continue to ask the next prior prompt statement until all 4 prior prompt statements are called. If all 4 prior prompt statements are normally answered, that is, none of them contain the negative answer character "I don't know", then it is considered that question q passes the prior judgment. It should be noted that the number of prior prompt statements in the prior prompt information group is not the more the better, because too many prompt statements will affect the accuracy of the prior judgment and also affect the processing efficiency. The default number of prompt statements can be 4.

[0028] It should be noted that during the process of reasoning about the factual knowledge questions in the factual knowledge question evaluation set in turn based on multiple prompt statements in the prior prompt information group through the linear verification method, it may also include: if the current reasoning answer contains a preset target character, then pause the current reasoning operation and directly determine that the pre-trained language model does not have the ability to answer the corresponding factual knowledge question in the factual knowledge question evaluation set. That is, during the process of reasoning about a single factual knowledge question in turn through multiple prompt statements (Prompt, used to prompt the text or instruction input by the user) in the prior prompt information group, if it is detected that the current reasoning answer contains a preset target character, such as "I don't know", then immediately stop the current reasoning operation, and at the same time determine that the pre-trained language model does not have the ability to answer the corresponding factual knowledge question q. That is, when the answer to any one of the prompt statements is "I don't know", it is considered that the model does not have the ability to answer this factual knowledge question. Immediately terminating the current reasoning operation when it is detected that the reasoning answer contains a preset target character can avoid subsequent invalid reasoning judgments, thereby improving the reasoning efficiency of the model.

[0029] In this embodiment, in order to more accurately and quickly reason about the factual knowledge questions in the factual knowledge question evaluation set, before reasoning, multiple prompt statements in the prior prompt information group can also be sorted. For example, the multiple prompt statements in the prompt information group are sorted according to the statement length of the prompt statements, the content complexity, the number of prompt verbs included, etc. Then, the sorted multiple prompt statements are used to reason about the factual knowledge questions in the factual knowledge question evaluation set, so as to obtain multiple reasoning answers for the corresponding factual knowledge questions. In addition, there can also be an association relationship between different prompt statements. For example, prompt statement a is a combination of prompt statement b and prompt statement c. Through the above diverse prompt statement design and sorting, the accuracy of the model's reasoning about factual knowledge questions can be further improved.

[0030] Step S13: If none of the multiple reasoning answers contain a preset target character, input the multiple reasoning answers and the preset posterior prompt information into the pre-trained language model to detect the authenticity of the multiple reasoning answers using the posterior prompt information, and obtain a detection result; the preset target character is a character that represents an abnormality in the reasoning answer.

[0031] In this embodiment, if none of the multiple reasoning answers corresponding to a certain factual knowledge question q contain a preset target character, and the preset target character is a character that represents an abnormality in the reasoning answer (such as "I don't know"), then the multiple reasoning answers and the preset posterior prompt information are input into the pre-trained language model to detect the authenticity of the multiple reasoning answers corresponding to a certain factual knowledge question q using the posterior prompt information, and obtain a detection result corresponding to the factual knowledge question q. For example, use the posterior prompt information (Is the following text fact true and correct? If not, please answer "Fact error") to detect the authenticity of the multiple reasoning answers corresponding to a certain factual knowledge question q. If the detection result contains the character "Fact error", it is determined that the multiple reasoning answers corresponding to the factual knowledge question q do not match the fact; if the detection result does not contain the character "Fact error", it is determined that the multiple reasoning answers corresponding to the factual knowledge question q match the fact.

[0032] In a specific embodiment, multiple inference answers and preset posterior prompt information are input into a pre-trained language model to detect the authenticity of the multiple inference answers using the posterior prompt information, and the detection result can specifically include: inputting the multiple inference answers and the preset first posterior prompt information into the pre-trained language model to integrate the multiple inference answers into a piece of text based on the first posterior prompt information, obtaining the integrated text; inputting the integrated text and the preset second posterior prompt information into the pre-trained language model to detect the authenticity of the integrated text based on the second posterior prompt information, obtaining the detection result. In this embodiment, in order to overcome the hallucination problem existing in the pre-trained language model, that is, it cannot accurately answer some factual knowledge questions, a posteriori check is further carried out on the basis of prior judgment, so as to prevent the model from giving wrong answers to factual knowledge questions due to hallucination. Specifically, referring to Figure 2 As shown, prior judgment specifically refers to the process of linearly reasoning about a certain factual knowledge question q input through a large language model and based on a preset group of prior prompt information. If there are 4 prompt statements in the group of prior prompt information, 4 corresponding inference answers can be obtained, which can be denoted as . If any one of the inference answers contains a specific character (such as "I don't know"), it is directly determined that the large language model does not have the ability to answer the question q. If none of the 4 inference answers contains a specific character (such as "I don't know"), it indicates that the prior judgment has been passed and the posteriori check can be carried out; referring to Figure 2 As shown, the posteriori check specifically refers to detecting the authenticity of multiple inference answers through a large language model and based on two preset posterior prompt information. Specifically, multiple inference answers can be integrated into a piece of text through a large language model and based on the first posterior prompt information (for example, "Please synthesize the following answers into a piece of text: {answer list}"), obtaining the integrated text A; then, the integrated text A and the second posterior prompt information (for example, Is the following text fact true and correct? If not, please answer "Fact error") are input into the large language model to detect the authenticity of the integrated text A to obtain the corresponding detection result. Through the two stages of prior judgment and posteriori check, this application can not only reason about factual knowledge questions, but also detect the authenticity of the inference answers obtained after reasoning, thus effectively preventing the large model from giving wrong answers to factual knowledge questions due to hallucination and preventing the large model from misjudging factual knowledge questions due to hallucination. In addition, synthesizing the multiple answers (i.e., multiple inference answers) of the prior judgment into a coherent piece of text and verifying the authenticity of the answer through the posterior prompt information can reduce the risk of misjudgment caused by the fragmentation of the answer (i.e., the inference answer).

[0033] Further, after detecting the authenticity of the integrated text based on the second posterior prompt information and obtaining the detection result, it may further include: if the detection result indicates that the integrated text does not conform to the facts, it is determined that the pre-trained language model does not have the ability to answer the corresponding factual knowledge questions in the factual knowledge question evaluation set. For example, refer to Figure 2 As shown, if the character "fact error" is included in the detection result, it is directly determined that the large language model does not have the ability to answer the corresponding factual knowledge question q.

[0034] Step S14: If the detection result indicates that multiple reasoning answers conform to the facts, it is determined that the pre-trained language model has the ability to answer the corresponding factual knowledge questions in the factual knowledge question evaluation set.

[0035] In this embodiment, if the detection result indicates that multiple reasoning answers conform to the facts, for example, the specific character "fact error" is not included in the detection result, it indicates that the posterior check passes. At this time, it can be determined that the large language model has the ability to answer the corresponding factual knowledge question q.

[0036] It can be seen that when evaluating the answering ability of the factual knowledge questions of the pre-trained language model in the embodiments of the present application, it is divided into two stages, specifically including the prior judgment and the posterior check. The first stage is to create a factual knowledge question evaluation set based on the historical knowledge graph of a certain field, and then use multiple prompt statements in the preset prior prompt information group to reason about each factual knowledge question in the factual knowledge question evaluation set in turn. The second stage is to input the multiple reasoning answers corresponding to each factual knowledge question obtained by reasoning and the preset posterior prompt information into the pre-trained language model, so as to detect the authenticity of the multiple reasoning answers. If the detection result indicates that the multiple reasoning answers conform to the facts, it is determined that the model has the ability to answer the corresponding factual knowledge questions. Through the above two stages (i.e., the prior judgment stage and the posterior check stage), and the prompt information preset in each stage, not only can the model be guided to judge whether it has the ability to answer the corresponding factual knowledge questions, but also through the posterior check operation, it can effectively prevent the model from misjudging the factual knowledge questions due to hallucinations. In addition, through the above method, without introducing additional model training calculation amount, the detection of the model's answering ability of factual knowledge can be realized, thereby improving the evaluation efficiency of the answering ability of factual knowledge questions and the model's cognitive ability of its own knowledge state.

[0037] The embodiments of the present application disclose a specific method for evaluating the answering ability of factual knowledge questions. Refer to Figure 3 As shown, the method includes: Step S21: Collect the historical knowledge graph of the target field, and sample the subgraphs in the historical knowledge graph based on different sampling rules to obtain multiple subgraph sets.

[0038] Step S22: Merge multiple sub - graph sets to obtain an evaluation sub - graph set, and input the evaluation sub - graph set into a pre - created factual knowledge question generation template to construct a factual knowledge question evaluation set covering the evaluation sub - graph set.

[0039] Step S23: Input the factual knowledge question evaluation set and a preset prior prompt information group into a pre - trained language model, and sequentially reason about the factual knowledge questions in the factual knowledge question evaluation set based on multiple prompt statements in the prior prompt information group through linear verification to obtain multiple reasoning answers corresponding to the factual knowledge questions.

[0040] Step S24: If none of the multiple reasoning answers contain a preset target character, input the multiple reasoning answers and a preset posterior prompt information into the pre - trained language model to detect the authenticity of the multiple reasoning answers using the posterior prompt information, and obtain a detection result; the preset target character is a character indicating that there is an abnormality in the reasoning answer.

[0041] Step S25: If the detection result indicates that the multiple reasoning answers are consistent with the facts, determine that the pre - trained language model has the ability to answer the corresponding factual knowledge questions in the factual knowledge question evaluation set.

[0042] Step S26: Based on the ability of the pre - trained language model to answer the factual knowledge questions in the factual knowledge question evaluation set and a preset composite evaluation index, statistically analyze the accuracy and knowledge coverage of the answers of the pre - trained language model to obtain the model answer correct rate and the model knowledge coverage rate.

[0043] In this embodiment, after sequentially reasoning and judging each factual knowledge question in the factual knowledge question evaluation set and detecting the authenticity of the reasoning answers, the ability of the pre - trained language model to answer factual knowledge questions in the target domain can be comprehensively evaluated based on the multiple reasoning answers (i.e., the results of prior judgment) corresponding to all factual knowledge questions in the factual knowledge question evaluation set and the detection results (i.e., the results of posterior inspection). Specifically, based on the ability of the pre - trained language model to answer each factual knowledge question q in the factual knowledge question evaluation set and a preset composite evaluation index (including the correct rate and the coverage rate), the accuracy and knowledge coverage of the answers of the pre - trained language model can be statistically analyzed respectively to obtain the corresponding model answer correct rate (Accuracy) and the model knowledge coverage rate (Coverage).

[0044] Specifically, based on the ability of the pre-trained language model to answer factual knowledge questions in the factual knowledge question evaluation set and the preset composite evaluation indicators, the accuracy and knowledge coverage of the answers of the pre-trained language model are statistically obtained. The model answer correct rate and the model knowledge coverage rate can include: screening the first type of factual knowledge questions from the factual knowledge question evaluation set; the first type of factual knowledge questions are factual knowledge questions that do not contain the preset target character in the corresponding multiple inference answers; screening the second type of factual knowledge questions from the first type of factual knowledge questions; the second type of factual knowledge questions are the first type of factual knowledge questions whose detection results indicate that the corresponding multiple inference answers are consistent with the facts; calculating the ratio of the number of the second type of factual knowledge questions to the total number of factual knowledge questions in the factual knowledge question evaluation set to obtain the model answer correct rate of the pre-trained language model; counting the number of target nodes in the evaluation subgraph set; the target nodes are nodes that do not contain the preset target character in the corresponding multiple inference answers; calculating the ratio of the number of target nodes to the total number of nodes in the evaluation subgraph set to obtain the model knowledge coverage rate of the pre-trained language model. In this embodiment, the first type of factual knowledge questions can be screened from the factual knowledge question evaluation set first. This question is specifically a factual knowledge question that does not contain the preset target character (such as "I don't know") in the corresponding multiple inference answers, that is Figure 2 the factual knowledge question q judged by prior knowledge in Figure 2 ; then, the second type of factual knowledge questions are screened from the first type of factual knowledge questions. The second type of factual knowledge questions are the first type of factual knowledge questions whose detection results indicate that the corresponding multiple inference answers are consistent with the facts, that is ; In the formula, represents the number of factual knowledge questions judged by prior knowledge and checked by posterior knowledge, represents the total number of all questions in the factual knowledge question evaluation set. Through the model answer correct rate (Accuracy) indicator, the answer accuracy of the model to known questions can be reflected. The higher the value, the stronger the knowledge reliability of the model.

[0045] Next, count the evaluation subgraph set The number of nodes of the target nodes, where the target nodes are the nodes that do not contain a preset target character (such as "I don't know") in the corresponding multiple inference answers, and calculate the ratio of the number of target nodes to the total number of all nodes in the evaluation subgraph set to obtain the model knowledge coverage rate (Coverage). The specific calculation formula is: ; In the formula, represents the number of nodes that do not contain a preset target character in the corresponding multiple inference answers, that is, the number of nodes in the evaluation subgraph where the questions are correctly answered; represents the evaluation subgraph the total number of all nodes in.

[0046] Step S27: Based on the model answer correct rate and the model knowledge coverage rate, comprehensively evaluate the ability of the pre-trained language model to answer factual knowledge questions in the target domain to obtain a performance evaluation result.

[0047] In this embodiment, after calculating the answer correct rate and knowledge coverage rate of the pre-trained language model, the ability of the pre-trained language model to answer factual knowledge in the target domain can be directly comprehensively evaluated based on the calculated model answer correct rate and model knowledge coverage rate to obtain the corresponding performance evaluation result.

[0048] Specifically, based on the model answer correct rate and the model knowledge coverage rate, comprehensively evaluate the ability of the pre-trained language model to answer factual knowledge questions in the target domain to obtain a performance evaluation result, which may include: calculating the product of the model answer correct rate and the first weight coefficient to obtain a first calculation result, and calculating the product of the model knowledge coverage rate and the second weight coefficient to obtain a second calculation result; the sum of the first weight coefficient and the second weight coefficient is 1; calculate the sum of the first calculation result and the second calculation result to obtain the performance evaluation result of the pre-trained language model. In this embodiment, the calculation formula of the performance evaluation result can be specifically expressed as: ; In the formula, represents the first weight coefficient, that is, the weight coefficient of the correct rate, and the value of this coefficient can reflect the degree of emphasis on accuracy; represents the second weight coefficient, that is, the weight coefficient of the coverage rate, and the value of this coefficient can reflect the degree of attention to the comprehensiveness of knowledge. In a specific implementation manner, and , and satisfy , representing the weight normalization constraint to ensure that the comprehensive score (i.e., the performance evaluation result) is in the range of [0,1].

[0049] The comprehensive performance of the model is evaluated by the correct answer rate of the model, the knowledge coverage rate of the model, and the corresponding weight coefficients. This can not only balance accuracy and comprehensiveness, but also adjust the value according to requirements. For example, in the medical field, the value can be increased to emphasize the knowledge coverage ability.

[0050] Among them, for the more specific processing procedures of the above steps S21 to S25, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.

[0051] It can be seen that the embodiment of the present application evaluates the answer ability of the pre-trained language model for factual knowledge questions through a two-stage strategy of prior judgment (verifying whether the model "knows") and posterior inspection (verifying the authenticity of the answer). It can accurately judge whether the answer is correct, solve the misjudgment problem caused by hallucinations of the model, and at the same time reduce the risk of misjudgment caused by hallucinations of the model. In addition, a set of prompt statements are introduced in the prior judgment stage of the embodiment of the present application to guide the model to judge whether it has the ability to answer factual knowledge questions. After the model answers, a posterior inspection operation is added to prevent the model from misjudging factual knowledge due to hallucinations. Through such a strategy, the detection of the factual knowledge ability of the model can be realized without introducing the computational complexity of additional model training, providing a basis for applications such as preventing large model hallucinations and strengthening the model with external knowledge. In addition, the embodiment of the present application can quantify the comprehensiveness of the model's knowledge mastery through a composite evaluation index including the correct rate (used to count the proportion of questions verified in two stages and measure the accuracy of the model's answers) and the coverage rate (used to count the proportion of knowledge nodes correctly answered and measure the knowledge coverage range of the model), and solves the limitation problem of traditional single indicators.

[0052] For example, collect a knowledge graph in the scientific field (such as Wikidata), then sample the "physicist" sub-graph in the knowledge graph (such as Wikidata), and input the set of evaluated sub-graphs obtained after sampling into a pre-created template for generating factual knowledge questions, so as to construct a factual knowledge question evaluation set covering the entire set of evaluated sub-graphs. This evaluation set includes multiple factual knowledge questions, such as high-frequency knowledge questions: "What is the nationality of Scientist M?" (degree = 120), and long-tail knowledge questions: "____ experiment first measured the electric charge of electrons" (entropy value = 3.2). Then, conduct a priori judgment and posterior inspection on each question in the evaluation set. The specific process is as follows: the question "____ experiment first measured the electric charge of electrons" → entity inspection passed → 4 Prompts verified passed → comprehensive reasoning answer "oil drop experiment" → posterior inspection passed. Further, evaluate the entire factual knowledge question evaluation set. The specific process is as follows: Accuracy = 87% → Coverage = 80% → Score = 0.7×87% + 0.3×80% = 84.9%, that is, the comprehensive evaluation score of the model's ability to answer factual knowledge in the scientific field is 84.9%.

[0053] The embodiments of the present application also provide an apparatus for evaluating the answering ability of factual knowledge questions. For the description of the features in the corresponding embodiments of the apparatus for evaluating the answering ability, reference can be made to the relevant descriptions in the corresponding embodiments of the method for evaluating the answering ability, which will not be elaborated here one by one.

[0054] Furthermore, the embodiments of the present application also disclose an electronic device Figure 4 is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be regarded as any limitation on the scope of use of the present application.

[0055] Figure 4 is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the method for evaluating the answering ability of factual knowledge questions disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0056] In this embodiment, the power supply 23 is used to provide operating voltages for the various hardware devices on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and specific limitations thereof are not provided herein; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type thereof can be selected according to specific application requirements, and specific limitations thereof are not provided herein.

[0057] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0058] Among them, the operating system 221 is used to manage and control the various hardware devices and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of evaluating the answering ability of factual knowledge questions executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs capable of performing other specific tasks.

[0059] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the method for evaluating the answering ability of factual knowledge questions disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.

[0060] Furthermore, the embodiments of this application also disclose a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method for evaluating the answering ability of factual knowledge questions disclosed as above are implemented.

[0061] The various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0062] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0063] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0064] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0065] The above has introduced in detail the method, product, device and medium for evaluating the ability to answer factual knowledge questions provided by this application. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for evaluating the answering ability of factual knowledge questions, characterized in that, Including: Collecting a historical knowledge graph of the target domain and sampling sub-graphs in the historical knowledge graph to generate a factual knowledge question evaluation set for the target domain; Inputting the factual knowledge question evaluation set and a preset prior prompt information group into a pre-trained language model, and sequentially reasoning about the factual knowledge questions in the factual knowledge question evaluation set based on multiple prompt statements in the prior prompt information group through a linear verification method to obtain multiple reasoning answers corresponding to the factual knowledge questions; If none of the multiple reasoning answers contain a preset target character, inputting the multiple reasoning answers and preset posterior prompt information into the pre-trained language model to detect the authenticity of the multiple reasoning answers using the posterior prompt information and obtain a detection result; The preset target character is a character indicating that the reasoning answer is abnormal; If the detection result indicates that the multiple reasoning answers are consistent with the facts, it is determined that the pre-trained language model has the ability to answer the corresponding factual knowledge questions in the factual knowledge question evaluation set.

2. The method for evaluating the answering ability of factual knowledge questions according to claim 1, characterized in that, The sampling of sub-graphs in the historical knowledge graph to generate a factual knowledge question evaluation set for the target domain includes: Sampling sub-graphs in the historical knowledge graph based on different sampling rules to obtain multiple sub-graph sets; Merging the multiple sub-graph sets to obtain an evaluation sub-graph set, and inputting the evaluation sub-graph set into a pre-created factual knowledge question generation template to construct a factual knowledge question evaluation set covering the evaluation sub-graph set.

3. The method for evaluating the answering ability of factual knowledge questions according to claim 2, characterized in that, The sampling of sub-graphs in the historical knowledge graph based on different sampling rules to obtain multiple sub-graph sets includes: Sampling sub-graphs in the historical knowledge graph based on a high-frequency knowledge sampling strategy and a long-tail knowledge sampling strategy to obtain a high-frequency sub-graph set and a long-tail sub-graph set; Among them, the high-frequency knowledge sampling strategy is a sampling strategy that takes the nodes with the top preset number of node degrees among the nodes in the knowledge graph as high-frequency knowledge nodes; the long-tail knowledge sampling strategy is a sampling strategy that takes the nodes in the sub-graphs of the knowledge graph with a sub-graph entropy value greater than the preset entropy value as long-tail knowledge nodes.

4. The method for evaluating the answering ability of factual knowledge questions according to claim 3, characterized in that, The merging of the multiple sub-graph sets to obtain an evaluation sub-graph set, and inputting the evaluation sub-graph set into a pre-created factual knowledge question generation template to construct a factual knowledge question evaluation set covering the evaluation sub-graph set includes: Merging the high-frequency sub-graph set and the long-tail sub-graph set to obtain an evaluation sub-graph set, and inputting the evaluation sub-graph set into a pre-created factual knowledge question generation template to construct a factual knowledge question evaluation set covering the evaluation sub-graph set.

5. The method for evaluating the answering ability of factual knowledge questions according to claim 4, characterized in that, The types of the factual knowledge question generation template include direct question type, triple mask type, and path query type; Among them, the direct question type is the type of directly generating questions based on triples in the knowledge graph; the triple masking type is the type of generating fill-in-the-blank questions by randomly masking the head entity or the tail entity; the path query type is the type of extracting relationship paths with lengths less than or equal to a preset length and generating multi-hop reasoning questions based on the relationship paths.

6. The method for evaluating the answering ability of factual knowledge questions according to claim 5, characterized in that, The process of inputting the evaluation sub-graph set into a pre-created factual knowledge question generation template to construct a factual knowledge question evaluation set covering the evaluation sub-graph set includes: Inputting the evaluation sub-graph set into a pre-created factual knowledge question generation template including the direct question type, the triple masking type, and the path query type, and generating questions corresponding to the direct question type, the triple masking type, and the path query type in corresponding quantities according to a preset generation ratio, to obtain a factual knowledge question evaluation set covering the evaluation sub-graph set; Among them, the number of questions generated by each sub-graph in the evaluation sub-graph set is determined based on the number of nodes of the corresponding sub-graph, and the ratio of the number of questions generated by the high-frequency sub-graph set to the number of questions generated by the long-tail sub-graph set is a preset value.

7. The method for evaluating the answering ability of factual knowledge questions according to claim 1, characterized in that, In the process of reasoning about the factual knowledge questions in the factual knowledge question evaluation set in turn based on the multiple prompt statements in the prior prompt information group by means of linear verification, it further includes: If the preset target character is included in the current reasoning answer, suspend the current reasoning operation and directly determine that the pre-trained language model does not have the ability to answer the corresponding factual knowledge questions in the factual knowledge question evaluation set.

8. The method for evaluating the answering ability of factual knowledge questions according to claim 2, characterized in that, After determining that the pre-trained language model has the ability to answer the corresponding factual knowledge questions in the factual knowledge question evaluation set, it further includes: Based on the ability of the pre-trained language model to answer the factual knowledge questions in the factual knowledge question evaluation set and a preset composite evaluation index, count the accuracy and knowledge coverage of the answers of the pre-trained language model to obtain the model answer correct rate and the model knowledge coverage rate; Based on the model answer correct rate and the model knowledge coverage rate, comprehensively evaluate the ability of the pre-trained language model to answer the factual knowledge questions in the target domain to obtain a performance evaluation result.

9. The method for evaluating the answering ability of factual knowledge questions according to claim 8, characterized in that, Based on the ability of the pre-trained language model to answer the factual knowledge questions in the factual knowledge question evaluation set and a preset composite evaluation index, count the accuracy and knowledge coverage of the answers of the pre-trained language model to obtain the model answer correct rate and the model knowledge coverage rate, including: Screen the first type of factual knowledge questions from the factual knowledge question evaluation set; the first type of factual knowledge questions are the factual knowledge questions in which the preset target character is not included in the corresponding multiple reasoning answers; Screen the second type of factual knowledge questions from the first type of factual knowledge questions; the second type of factual knowledge questions are the first type of factual knowledge questions in which the detection result indicates that the corresponding multiple reasoning answers are consistent with the facts; Calculate the ratio of the number of questions in the second type of factual knowledge questions to the total number of factual knowledge questions in the factual knowledge question evaluation set to obtain the model answer correct rate of the pre-trained language model; Count the number of nodes of the target nodes in the evaluation sub-graph set; the target nodes are the nodes that do not contain the preset target character in the corresponding multiple reasoning answers; Calculate the ratio of the number of nodes of the target nodes to the total number of nodes in the evaluation sub-graph set to obtain the model knowledge coverage rate of the pre-trained language model.

10. The method for evaluating the ability to answer factual knowledge questions according to claim 8, characterized in that, Based on the model answer correct rate and the model knowledge coverage rate, comprehensively evaluate the ability of the pre-trained language model to answer factual knowledge questions in the target domain to obtain a performance evaluation result, including: Calculate the product of the model answer correct rate and the first weight coefficient to obtain a first calculation result, and calculate the product of the model knowledge coverage rate and the second weight coefficient to obtain a second calculation result; the sum of the first weight coefficient and the second weight coefficient is 1; Calculate the sum of the first calculation result and the second calculation result to obtain the performance evaluation result of the pre-trained language model.

11. The method for evaluating the ability to answer factual knowledge questions according to any one of claims 1 to 10, characterized in that, Input the multiple reasoning answers and the preset posterior prompt information into the pre-trained language model to detect the authenticity of the multiple reasoning answers by using the posterior prompt information, and the obtained detection result includes: Input the multiple reasoning answers and the preset first posterior prompt information into the pre-trained language model to integrate the multiple reasoning answers into a piece of text based on the first posterior prompt information to obtain an integrated text; Input the integrated text and the preset second posterior prompt information into the pre-trained language model to detect the authenticity of the integrated text based on the second posterior prompt information to obtain a detection result.

12. The method for evaluating the ability to answer factual knowledge questions according to claim 11, characterized in that, After detecting the authenticity of the integrated text based on the second posterior prompt information to obtain a detection result, it further includes: If the detection result indicates that the integrated text does not conform to the facts, it is determined that the pre-trained language model does not have the ability to answer the corresponding factual knowledge questions in the factual knowledge question evaluation set.

13. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the method for evaluating the answer ability of factual knowledge questions according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the method for evaluating the answer ability of factual knowledge questions according to any one of claims 1 to 12.

15. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method for evaluating the answer ability of factual knowledge questions according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Data processing method of dialogue system, electronic equipment and readable storage medium

    CN114090757A

  • Single-sample time sequence knowledge graph extrapolation calculation method based on historical trend

    CN117575023A

  • Insurance knowledge question-answer pair acquisition method and device, equipment and storage medium

    CN117709359A

  • Graph attention mechanism feature fusion method of factual knowledge question-answering system

    CN117892256A

  • Knowledge graph multi-order reasoning enhanced question and answer method and computer readable medium

    CN118296115A

Cited By

  • Knowledge cross validation question and answer method and system for reducing illusion of large language model

    CN120910218A

  • Confidence-driven adaptive interview evaluation method and device, medium and equipment

    CN122470761A