A Two-Stage Few-Shot Automatic Fact-Checking Method, Electronic Device, and Storage Medium
Through the two-stage automatic fact verification method of few samples, the BM25, MonoT5 and Flan-T5 models are used for evidence retrieval and declaration verification, which solves the problem of quickly discovering false information and achieves efficient and low-cost automatic fact verification.
Patent Information
- Application Number
- CN202411828667.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-12-12
AI Technical Summary
In the era of information explosion, it has become a problem to quickly and accurately discover false information. Manual verification is time-consuming and resource-intensive, and the development of automatic fact verification technology is urgently needed.
A two-stage automatic fact verification method for few samples is adopted, including evidence retrieval and declaration verification, document and sentence selection is used using BM25 and MonoT5 models, and fine-tuning is combined with Flan-T5 language model to achieve few samples training.
Without a large amount of training data, the accuracy and efficiency of information verification are improved, the cost of data collection and computing resources is reduced, and efficient automatic fact verification is achieved.
Smart Images

Figure CN119691152B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information processing, and particularly relates to a two-stage few-shot automatic fact-checking method, an electronic device, and a storage medium. Background Art
[0002] In the current era of rapid development of the Internet, people are surrounded by a large amount of information every day, which comes from various sources, including social media, news websites, blogs, and other platforms. Whether as ordinary people or journalists, it has become increasingly difficult to judge the authenticity of information in the vast amount of information. Therefore, in this era of "information explosion", how to quickly and accurately discover false information has become an urgent problem to be solved.
[0003] With the popularization of social media and the rapid growth of digital content, fake news and misleading information have also increased accordingly. This type of information may mislead the public, intensify social differences, and have a negative impact on public trust. In addition, although many news agencies and independent groups are committed to fact-checking, manual fact-checking is a time-consuming and resource-intensive process, and it is difficult to keep up with the rapid flow of information in real time, especially in the Internet era when new topics and hotspots emerge continuously.
[0004] Therefore, it is very necessary to develop automatic fact-checking technology. First of all, automatic fact-checking technology can quickly process a large amount of information, provide a preliminary authenticity assessment, and assist human fact-checkers to conduct fact-checking more efficiently. Through automated tools, misleading information can be identified and corrected faster, thus protecting the public from being misled by fake news and enhancing trust in the media and other information sources. Secondly, accurate and timely information is crucial for various decision-makers, such as policymakers, business leaders, and the general public. Automatic fact-checking can ensure that decisions are based on true and accurate information. In addition, automatic fact-checking is a cutting-edge research topic in the field of natural language processing, and its research will further promote the development of related technologies. Summary of the Invention
[0005] The problem to be solved by the present invention is to improve the accuracy and efficiency of information fact-checking, and a two-stage few-shot automatic fact-checking method, an electronic device, and a storage medium are proposed.
[0006] To achieve the above object, the present invention is realized through the following technical solutions:
[0007] A two-stage few-shot automatic fact-checking method includes the following steps:
[0008] S1. Select a publicly available fact-checking dataset to construct a training set and a test set, which include claims to be verified and a knowledge base;
[0009] S2. Use BM25 to build an index for the documents in the knowledge base as a document retrieval model. Take the claims in the test set in step 1 as queries, and retrieve the top several documents most relevant to the claims in step S1 as candidate documents;
[0010] S3. Build a sentence selection model. Input each sentence in the candidate documents obtained in step S2 and the claims in step S1 into the prompt template prompt to form the input of the sentence selection model, and input it into the sentence selection model to output the relevance score of the sentence;
[0011] S4. Based on the relevance scores of the sentences obtained in step S3, select the top several sentences as the evidence sentences for claim verification;
[0012] S5. Build a pre-trained language model and fine-tune the pre-trained language model using the few-shot training set extracted from the training set;
[0013] S6. Input the evidence sentences for claim verification obtained in step S4 and the claim to be verified into the prompt template prompt to form the input of the pre-trained language model, and input it into the fine-tuned pre-trained language model to obtain a natural language output, then map it to classification labels, and calculate the prediction score based on the generation probability of the output sequence to obtain the final prediction result.
[0014] Further, the public fact verification dataset described in step S1 includes one or a combination of several of the FEVER dataset, the SciFact dataset, and the VitaminC dataset. The FEVER dataset makes artificially generated claims by modifying the content in Wikipedia and provides standard evidence for the claims with support and refutation labels; the SciFact dataset consists of scientific claims formed by experts re-writing sentences in biomedical literature, and the VitaminC dataset is created by using the revised version of Wikipedia to construct a test set.
[0015] Further, the specific implementation method of step S2 includes the following steps:
[0016] S2.1. Build an evidence retrieval model as the probabilistic retrieval model BM25. Calculate the frequency of the query term appearing in the document and the vocabulary distribution in the document through statistical methods to generate the query term q i The relevance score score(q, d) in document d, and the expression is:
[0017]
[0018] where q represents the query term, d represents the document, IDF(q i ) represents the inverse document frequency of the query term q i and f(qi , d) represents the query term q i The frequency of occurrence of q in document d, |d| and avgdl represent the document length and the average length of all documents respectively, and k1 and b are the first adjustable parameter and the second adjustable parameter respectively;
[0019] S2.2. Use Elastic Search to build BM25 indexes for SciFact and FEVER respectively. Based on the evidence retrieval model constructed in step S2.1, retrieve the test set obtained in step S1 based on the claim, and evaluate the retrieval results. The evaluation criteria used are ndcg@n and recall@n, and the top 100 documents most relevant to the claim in step S1 are obtained as candidate documents.
[0020] Furthermore, the specific implementation method of step S3 includes the following steps:
[0021] S3.1. Build a sentence selection model, and use the checkpoint released by the MonoT5 model for the sentence selection task. MonoT5 is a retrieval model fine-tuned using the MS MARCO paragraph dataset, and use the checkpoint released by the model for the sentence selection task;
[0022] The sentence selection model generates target tokens of 'True' or 'False' according to whether the document is relevant to the claim. In the ranking task, aggregate the output results of the sentence selection model, calculate the probability of assigning the token 'True' as the relevance score, and generate the relevance score of the sentence. The expression is:
[0023]
[0024] where relevant = 1 indicates that the claim and the document are relevant, logit('True') represents the correct assignment probability of the target token, and logit('False') represents the incorrect assignment probability of the target token;
[0025] S3.2. Split the candidate documents obtained in step S2 into sentences, and then combine each sentence in the candidate documents obtained in step S2 with the claim in step S1 and input them into the prompt template prompt. The expression of the prompt template prompt is:
[0026] Query: [Q] Document: [D] Relevant:
[0027] where Q is the claim in the dataset and D is the document sentence to be evaluated;
[0028] S3.3. Input the combination obtained in step S3.2 into the sentence selection model for re-ranking, and output the relevance score of the sentence.
[0029] Further, in step S4, the top five ranked sentences are selected as the evidence sentences for claim verification.
[0030] Further, the specific implementation method of step S5 includes the following steps:
[0031] S5.1. Select Flan-T5 as the base model of the pre-trained language model;
[0032] S5.2. Fine-tune the base model of the pre-trained language model using the LoRA fine-tuning method. During training, the rank of the matrix is specified as 4, and LoRA is applied to the (q|k|v|o|w) layers of the SelfAttention, EncDecAttention, and DenseReluDense modules of the original model to obtain the pre-trained language model.
[0033] Further, the specific implementation method of step S6 includes the following steps:
[0034] S6.1. Set the input of the fact-checking task to consist of the claim C to be verified i and the evidence sentences e for claim verification obtained in step S4 i . Substitute (C i , e i ) into the prompt template to obtain the output x i mapped to the label y i . An example of the prompt template is:
[0035] Suppose{e i}, can we infer that{C i}?;
[0036] S6.2. Based on the model input being x i , generate the probability p i of a target sequence y θ : i , x i ):
[0037]
[0038] where p θ (t j |x i , t <j ) is the probability that the model generates t i given the input sequence x <j and the previously generated sequence t j ;
[0039] Due to the correspondence between the output sequence and the classification label, the predicted score of the model for y i is defined as the logarithm length normalization of the probability of generating the target sequence, resulting in:
[0040]
[0041] Use the predicted scores to rank all categories, and obtain the category with the first rank as the final category prediction.
[0042] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the two-stage few-shot automatic fact-checking method are implemented.
[0043] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the two-stage few-shot automatic fact-checking method is implemented.
[0044] Advantages of the present invention:
[0045] The two-stage few-shot automatic fact-checking method described in the present invention includes few-shot evidence retrieval and few-shot claim verification. In claim verification, using the evidence sentences (perfect evidence) given in the dataset, under the condition of no more than 50 groups of training data, the accuracy rate reaches the SOTA (current highest level) of the fully supervised model. In the few-shot evidence sentence exact retrieval task, a retrieval method combining BM25 and the Seq2Seq language model is adopted, and the obtained effect is slightly inferior to that reported by the fully supervised model.
[0046] The two-stage few-shot automatic fact-checking method described in the present invention adopts few-shot settings in both stages of the fact-checking task. This setting method enables the entire fact-checking link to achieve good results without too much training data. This has important significance in practical applications and can greatly reduce the cost and workload of data collection and processing.
[0047] The two-stage few-shot automatic fact-checking method described in the present invention uses a relatively small number of model parameters. A smaller number of parameters means lower training costs and computational resource requirements. Compared with other models of the same scale, the method of the present invention shows the best effect on the FEVER dataset. This fully proves the effectiveness and superiority of the method and provides an efficient and low-cost solution for the field of fact-checking. Description of the Drawings
[0048] Figure 1 is a flowchart of the two-stage few-shot automatic fact-checking method described in the present invention;
[0049] Figure 2The structural block diagram of a two-stage few-shot automatic fact-checking method according to the present invention;
[0050] Figure 3 The effect comparison diagram of the model of the present invention and ProToCo-3B, where (a) is the comparison of the SciFact dataset, (b) is the comparison of the FEVER dataset, and (c) is the comparison of the VitaminC dataset. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. It should be understood that the specific implementation manners described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific implementation manners described are only a part of the implementation manners of the present invention, rather than all the specific implementation manners. The components of the specific implementation manners of the present invention usually described and shown in the accompanying drawings here can be arranged and designed in various different configurations, and the present invention can also have other implementation manners.
[0052] Therefore, the detailed description of the specific implementation manners of the present invention provided in the accompanying drawings below is not intended to limit the scope of the claimed present invention, but only represents the selected specific implementation manners of the present invention. Based on the specific implementation manners of the present invention, all other specific implementation manners obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0053] In order to further understand the content, features and effects of the present invention, the following specific implementation manners are exemplified and combined with the attached Figure 1 - Attachment Figure 3 The details are as follows:
[0054] Example 1:
[0055] A two-stage few-shot automatic fact-checking method includes the following steps:
[0056] S1. Select a publicly available fact-checking dataset to construct a training set and a test set, which includes claims to be verified and a knowledge base;
[0057] Further, the publicly available fact-checking dataset described in step S1 includes one or a combination of several of the FEVER dataset, the SciFact dataset, and the VitaminC dataset. The FEVER dataset makes artificially generated claims by modifying the content in Wikipedia and provides standard evidence for the claims with support and refutation labels; the SciFact dataset consists of scientific claims formed by experts re-writing sentences in biomedical literature, and the VitaminC dataset is created by using the revised version of Wikipedia to construct a test set;
[0058] The relevant statistical information of the dataset is shown in Table 1:
[0059] Table 1 Statistical Information of the Collected Dataset
[0060]
[0061] Since the dataset only provides perfect evidence for the claim of determining the label, some sentences are randomly selected as evidence for the claim of the NEI label.
[0062] S2. Use BM25 to build an index for the documents in the knowledge base as a document retrieval model. Take the claims in the test set in step 1 as queries, and retrieve the top several documents most relevant to the claims in step S1 as candidate documents;
[0063] Evidence retrieval is a necessary part of the fact-checking pipeline, aiming to find evidence related to the claim in a large number of documents. The evidence often needs to be relatively fine-grained fragments such as sentences. After research, both FEVER and Scifact are included in a retrieval benchmark called BEIR (BenchMarking Information Retrieval) to evaluate the retrieval system, and both provide sentence-level relevant evidence annotation. Therefore, this embodiment is carried out on these two datasets and compared with other methods.
[0064] After research, the evidence retrieval of existing work is usually divided into two stages - document retrieval and sentence selection. Document retrieval selects documents related to the claim from a large number of documents, and sentence selection selects more precise evidence sentences from these documents. However, currently, the open-source evidence retrieval methods basically do not meet the few-shot and zero-shot conditions, that is, the retrieval model is trained using all the training data. Therefore, this embodiment studies how to retrieve possible evidence without using a large amount of training data to form a few-shot fact-checking pipeline.
[0065] Furthermore, the specific implementation method of step S2 includes the following steps:
[0066] S2.1. Build the evidence retrieval model as the probabilistic retrieval model BM25. Calculate the frequency of the query term appearing in the document and the vocabulary distribution in the document through statistical methods to generate the query term q i The relevance score score(q, d) of the query term q in the document d is expressed as:
[0067]
[0068] where q represents the query term, d represents the document, and IDF(q i ) represents the query term q iInverse document frequency of, f(q i , d) represents the query term q i The frequency of occurrence in document d, |d| and avgdl represent the document length and the average length of all documents respectively, and k1 and b are the first adjustable parameter and the second adjustable parameter respectively;
[0069] S2.2. Use Elastic Search to build BM25 indexes for SciFact and FEVER respectively. Based on the evidence retrieval model constructed in step S2.1, retrieve the test set obtained in step S1 based on the claim, and evaluate the retrieval results. The evaluation criteria used are ndcg@n and recall@n, and the top 100 documents most relevant to the claim in step S1 are used as candidate documents.
[0070] The results are shown in Table 2:
[0071] Table 2 Document retrieval results using BM25
[0072]
[0073] It can be seen that among the top 100 documents in the retrieval results, the recall rate reaches a relatively high level and can be used as the input document set for the next step of sentence selection.
[0074] S3. Build a sentence selection model. Combine each sentence in the candidate documents obtained in step S2 with the claim in step S1 and input them into the prompt template prompt as the input of the sentence selection model, and then input them into the sentence selection model to output the relevance score of the sentence;
[0075] Currently, the Transformer model pre-trained with language modeling objectives has been proven to be very effective in various classification and sequence tagging tasks in NLP. Using a complex retrieval model to re-rank the previous retrieval results has also become the standard architecture of all current Transformer-based document retrieval methods. This embodiment believes that sentence selection can also be regarded as a re-ranking task, so it is also applicable to these re-ranking retrieval models. Split the candidate documents described in step 2 into sentences, input them into the sentence selection model of this step for re-ranking, and select several pieces of evidence that best meet the requirements.
[0076] Furthermore, the specific implementation method of step S3 includes the following steps:
[0077] S3.1. Build a sentence selection model. Use the checkpoint released by the MonoT5 model for the sentence selection task. MonoT5 is a retrieval model fine-tuned using the MS MARCO passage dataset, and use the checkpoint released by the model for the sentence selection task;
[0078] The sentence selection model generates target tokens of 'True' or 'False' based on whether the document is relevant to the claim. In the ranking task, the output results of the sentence selection model are aggregated, and the probability of assigning the token 'True' is calculated as the relevance score to generate the relevance score of the sentence. The expression is as follows:
[0079]
[0080] Among them, relevant = 1 indicates that the claim and the document are relevant, logit('True') represents the correct assignment probability of the target token, and logit('False') represents the incorrect assignment probability of the target token;
[0081] Furthermore, MonoT5 is a retrieval model based on the T5 model (base, large, and 3B) and fine-tuned using the MS MARCO passage dataset; the method of using the above Seq2Seq model for the sentence ranking task is similar to claim verification, and both require constructing the input document and claim into the input through a prompt template, and then obtaining the similarity score using the output content.
[0082] S3.2. Split the candidate documents obtained in step S2 into sentences, and then combine each sentence in the candidate documents obtained in step S2 with the claim in step S1 into the prompt template prompt. The expression of the prompt template prompt is as follows:
[0083] Query: [Q] Document: [D] Relevant:
[0084] Among them, Q is the claim in the dataset, and D is the document sentence to be evaluated;
[0085] S3.3. Input the combination obtained in step S3.2 into the sentence selection model for re-ranking, and output the relevance score of the sentence.
[0086] S4. Based on the relevance scores of the sentences obtained in step S3, select the top several sentences as the evidence sentences for claim verification;
[0087] Furthermore, in step S4, select the top five sentences as the evidence sentences for claim verification;
[0088] S5. Construct a pre-trained language model and fine-tune the pre-trained language model using a few-shot training set extracted from the training set;
[0089] Furthermore, the specific implementation method of step S5 includes the following steps:
[0090] S5.1. Select Flan-T5 as the base model for the pre-trained language model;
[0091] Furthermore, models with the T5 (Text-to-Text Transfer Transformer) architecture were examined, and one of them was selected as the base model. T5 is an encoder-decoder Transformer model that is pre-trained by predicting masked target words on a large corpus of unlabeled text data. It treats all NLP tasks as text-to-text conversion problems and can be applied to a variety of downstream tasks. At the same time, models with the T5 architecture often come in various sizes to adapt to different computing resources and task requirements, and a version suitable for the application scenario can be selected.
[0092] T0 was created by fine-tuning T5 on a multi-task mixed dataset to achieve zero-shot generalization, that is, the ability to perform tasks without any additional gradient-based training. Examples in the dataset used to train T0 are prompted by applying prompt templates from the Public Pool of Prompts, which convert each example in each dataset into the text-to-text format of the prompt.
[0093] Flan-T5 is a series of model checkpoints released by Google that perform instruction fine-tuning (referred to as Flan fine-tuning) on a collection of data sources using various instruction template types and also incorporate several chain-of-thought datasets. These checkpoints have strong few-shot and zero-shot learning capabilities as well as chain-of-thought reasoning capabilities, outperforming previous T5 series checkpoints.
[0094] S5.2. Fine-tune the base model of the pre-trained language model using the LoRA fine-tuning method. The rank of the matrix is specified as 4 during training. Apply LoRA to the (q|k|v|o|w) layers of the SelfAttention, EncDecAttention, and DenseReluDense modules of the original model to obtain the pre-trained language model;
[0095] Furthermore, the core idea of the method is to approximate the update of the original weight matrix using the product of two low-rank matrices. By setting the rank of the matrix, the amount of parameters to be updated can be controlled. Therefore, LoRA can significantly reduce the number of parameters to be trained, reduce the demand for computing resources, and improve the stability in training with a small amount of data. At the same time, it is also possible to control the specific parts of the model to which LoRA is applied, such as the query (Q), key (K), and value (V) matrices in the self-attention mechanism.
[0096] During training, the rank of the matrix is specified as 4, and LoRA is applied to the (q|k|v|o|w) layers of the SelfAttention, EncDecAttention, and DenseReluDense modules in the original model.
[0097] In terms of the loss function used, the cross-entropy loss is adopted. Using the cross-entropy loss function will train the model in the direction of increasing the probability of the correct target sequence y i ;
[0098]
[0099] where p θ (t j |x i ,t <j ) is the probability that the model generates t i under the condition of the given input sequence x <j and the previously generated sequence t j , and y i is the target sequence;
[0100] Since it is a classification task, a prediction score can be calculated for each category, and thus the softmax probability of the label y i can be calculated. Maximizing this probability also helps the model increase the prediction probability of the correct category:
[0101]
[0102] L cls = -log q θ (y i |x i )
[0103] where γ is the set of all target sequences, and p θ (y i ,x i ) represents the probability that the model generates the target sequence y i under the condition of the given input sequence x <j and the previously generated sequence t i . exp represents the natural exponential function, and q θ (y i |x i ) represents the softmax probability of y i ;
[0104] Maximizing the probability of the correct class by the model is equivalent to minimizing the probability of the incorrect class. Therefore, this paper introduces the Unlikelihood loss, which enables the model to assign a lower probability to incorrect class sequences:
[0105]
[0106] where || represents the length;
[0107] In terms of training data selection, since fact-checking is a three-classification problem, for class balance, this paper refers to a training data containing a triple of {SUPPORT, RRFUTE, NEI} as a shot, and randomly selects training data with [1, 2, 4, 8, 16] shots from the dataset as the training set.
[0108] S6. Input the evidence sentence for claim verification obtained in step S4 and the claim to be verified into the prompt template prompt for combination as the input of the pre-trained language model, and input it into the fine-tuned pre-trained language model to obtain a natural language output, then map it to a classification label, calculate the prediction score based on the generation probability of the output sequence, and obtain the final prediction result.
[0109] Furthermore, the specific implementation method of step S6 includes the following steps:
[0110] S6.1. Set the input of the fact-checking task to consist of the claim C to be verified i and the evidence sentence e for claim verification obtained in step S4 i . Substitute (C i , e i ) into the prompt template to obtain the output x i and map it to the label y i . The expression of the prompt template is:
[0111] Suppose{e i}, can we infer that{C i}?;
[0112] Furthermore, the label y i has three classifications: support, refute, and insufficient information (SUPPORT, RRFUTE, NEI);
[0113] S6.2. Based on the model input being x i , generate the probability p i of a target sequence y θ : i , x i )
[0114]
[0115] Among them, p θ (t j |x i ,t <j ) is the probability that the model generates t under the condition of the given input sequence x i and the previously generated sequence t <j ; j
[0116] Due to the correspondence between the output sequence and the classification label, the prediction score of the model for y i is defined as the logarithm length normalization of the probability of generating the target sequence, obtaining:
[0117]
[0118] Use the prediction score to rank all categories, and obtain the category with the first rank as the final category prediction.
[0119] The experimental effects of this embodiment are as follows:
[0120] Use the training sets of [1, 2, 4, 8, 16] shot training data randomly selected from FEVER, SciFact, and VitaminC respectively. Apply LoRA fine-tuning on the (q|k|v|o|w) layers of the SelfAttention, EncDecAttention, and DenseReluDense modules of Flan-T5-large, and test on the validation sets of each dataset. Use the macro F1 value as the classification evaluation index. The experimental results are shown in Table 3.
[0121] Table 3 Macro-average F1 score of models fine-tuned with different sizes of training data on the validation set
[0122]
[0123] Furthermore, the results are also compared with existing few-shot fact-checking methods. The baseline model participating in the comparison is the 3B-sized checkpoint of ProToCo, which is also a fact-checking model based on the T5 architecture. It uses T0 as the backbone model and uses some consistency constraints to expand the training data for few-shot fine-tuning on relevant datasets. And the model size of this embodiment is 780M. The comparison results are as Figure 3 shown. The model of this embodiment exceeds the baseline method on each dataset with a smaller number of parameters, especially under the condition of extremely few samples.
[0124] In addition, on the FEVER dataset, there are many results of previous experiments where fully supervised models were trained using all the training data. The label accuracies of these models using perfect evidence were compared. The comparison results are shown in Table 4. The model fine-tuned with 4 shots, that is, 12 pieces of training data, has exceeded the accuracy of using all the training data.
[0125] Table 4 Comparison of the accuracies of fully supervised models and the models in this paper using FEVER perfect evidence
[0126]
[0127] Based on the results of BM25, the top 100 documents for each query were taken, and the sentences in them were scored and sorted using the base-sized monoT5. The retrieval results are shown in Table 5:
[0128] Table 5 Sentence retrieval results using monoT5
[0129]
[0130] Since the results of zero-shot evidence sentence retrieval are few, the comparable results are also very limited. On the FEVER dataset, the zero-shot evidence sentence retrieval results reported by LisT5 are recall@5 = 0.8539, but the T5 model size it uses is 3B, and the computing resources of this embodiment cannot run a model of this size. However, the gap can be compensated by expanding the candidate set size.
[0131] Based on the results of few-shot evidence retrieval, the top 5 scored evidence sentences were selected as evidence for each claim, and the models trained using few-shot training sets of different sizes were used as fact-checking models. The obtained fact-checking results are shown in Table 6. Compared with the excellent results presented using perfect evidence, the results using retrieved evidence have indeed declined in terms of effectiveness. The main reasons are as follows. The recall rate of the retrieved evidence in the first five pieces of evidence is only about 75%. This means that during the retrieval process, a certain proportion of relevant evidence fails to be recalled in a timely and accurate manner, thus affecting the final result performance. At the same time, there is also a part of noise mixed in the retrieved evidence, and these noises may come from various inaccurate matches or interferences from irrelevant information, further reducing the quality and effectiveness of the retrieved evidence. Moreover, it should be noted that there are usually only 1 to 2 pieces of perfect evidence annotated in the dataset. Due to the small quantity, these perfect evidences are often carefully selected and annotated, with high accuracy and reliability. In addition, the input length of the perfect evidence is short, which makes it more concise and clear, and also means a larger information density, capable of providing more powerful support for decision-making more efficiently. In contrast, the retrieved evidence has certain deficiencies in these aspects, resulting in its result effectiveness being inferior to that of perfect evidence.
[0132] Table 6 Few-shot Fact-Checking Results Based on Few-shot Evidence Retrieval
[0133]
[0134] Example 2:
[0135] An electronic device, characterized in that it includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of a two-stage few-shot automatic fact-checking method described in Example 1.
[0136] The computer device of the present invention may be a device including a processor and a memory, such as a single-chip microcomputer including a central processing unit. And, when the processor is used to execute the computer program stored in the memory, it implements the steps of the above-mentioned two-stage few-shot automatic fact-checking method.
[0137] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0138] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0139] Example 3:
[0140] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, it implements the two-stage few-shot automatic fact-checking method described in Embodiment 1.
[0141] The computer-readable storage medium of the present invention can be any form of storage medium readable by the processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned two-stage few-shot automatic fact-checking method can be implemented.
[0142] The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0143] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0144] Although the present application has been described above with reference to specific embodiments, various improvements can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the features in the specific embodiments disclosed in the present application can be combined with each other in any manner, and the non-exhaustive description of these combinations in this specification is only for the sake of saving space and resources. Therefore, the present application is not limited to the specific particular embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A two-stage few-shot automatic fact-checking method, characterized in that, It includes the following steps: S1. Select a publicly available fact-checking dataset to construct a training set and a test set, which contain claims to be verified and a knowledge base; S2. Use BM25 to construct an index for the documents in the knowledge base as a document retrieval model. Take the claims in the test set in step S1 as queries, and retrieve the top several documents most relevant to the claims in step S1 as candidate documents; S3. Construct a sentence selection model. Input each sentence in the candidate documents obtained in step S2 and the claims in step S1 into the prompt template prompt for combination as the input of the sentence selection model, and input it into the sentence selection model to output the relevance score of the sentence; The specific implementation method of step S3 includes the following steps: S3.
1. Construct a sentence selection model and use the checkpoint released by the MonoT5 model for the sentence selection task; S3.
2. Split the candidate documents obtained in step S2 into sentences, and then input each sentence in the candidate documents obtained in step S2 and the claims in step S1 into the prompt template prompt for combination. The expression of the prompt template prompt is: Query:[Q]Document:[D]Relevant: where Q is the claim in the dataset and D is the document sentence to be evaluated; S3.
3. Input the combination obtained in step S3.2 into the sentence selection model for re-ranking and output the relevance score of the sentence; S4. Based on the relevance scores of the sentences obtained in step S3, select the top several sentences as the evidence sentences for claim verification; S5. Construct a pre-trained language model and fine-tune the pre-trained language model using a few-shot training set extracted from the training set; S6. Input the evidence sentences for claim verification obtained in step S4 and the claims to be verified into the prompt template prompt for combination as the input of the pre-trained language model, and input it into the fine-tuned pre-trained language model to obtain a natural language output, then map it to a classification label, calculate the prediction score based on the generation probability of the output sequence, and obtain the final prediction result; The specific implementation method of step S6 includes the following steps: S6.
1. Set the input of the fact-checking task to be composed of the claim C to be verified i and the evidence sentence e for claim verification obtained in step S4 i . Substitute (C i , e i ) into the prompt template, and map the resulting output x i to the label y i . The prompt template is: Suppose{e i}, can we infer that{C i}?; S6.
2. Based on the model input x i , generate a target sequence y i with probability p θ (y i , x i ): where p θ (t j |x i ,t <j ) is the probability that the model generates t given the input sequence x i and the previously generated sequence t <j under the condition; j of generating t Due to the correspondence between the output sequence and the classification label, the predicted score of the model for y i is defined as the logarithm length normalization of the probability of generating the target sequence, resulting in: Rank all categories using the prediction score, and obtain the category ranked first as the final category prediction.
2. The two-stage few-shot automatic fact-checking method according to claim 1, wherein, The publicly available fact-checking dataset mentioned in step S1 includes one or a combination of the FEVER dataset, the SciFact dataset, and the VitaminC dataset. The FEVER dataset makes artificially generated claims by modifying the content in Wikipedia and provides standard evidence for the claims with support and refutation labels; the SciFact dataset consists of scientific claims formed by experts re-writing sentences in biomedical literature, and the VitaminC dataset is created by using the revised version of Wikipedia to construct a test set.
3. The two-stage few-shot automatic fact-checking method according to claim 2, wherein, The specific implementation method of step S2 includes the following steps: S2.
1. Construct the evidence retrieval model as the probabilistic retrieval model BM25, calculate the frequency of the query term in the document and the vocabulary distribution in the document through statistical methods, and generate the query term q i The relevance score score(q, d) of the query term q in the document d is expressed as: Among them, q represents the query term, d represents the document, IDF(q i ) represents the inverted document frequency of the query term q i , f(q i , d) represents the frequency of occurrence of the query term q i in the document d, |d| and avgdl respectively represent the document length and the average length of all documents, and k1 and b are the first adjustable parameter and the second adjustable parameter respectively; S2.
2. Build BM25 indexes for the SciFact and FEVER datasets using Elastic Search, retrieve the test set obtained in step S1 based on the claims using the evidence retrieval model constructed in step S2.1, and evaluate the retrieval results. The evaluation criteria used are ndcg@n and recall@n, and the top 100 documents most relevant to the claims in step S1 are obtained as candidate documents.
4. A two-stage few-shot automatic fact-checking method according to claim 3, characterized in that In step S3, the MonoT5 model is a retrieval model fine-tuned using the MS MARCO passage dataset, and the checkpoint released by the model is used for the sentence selection task; The sentence selection model generates target tokens of 'True' or 'False' according to whether the document is relevant to the claim. In the ranking task, the output results of the sentence selection model are aggregated, and the probability of assigning the token 'True' is calculated as the relevance score to generate the relevance score of the sentence. The expression is: where relevant = 1 indicates that the claim and the document are relevant, logit('True') represents the correct assignment probability of the target token, and logit('False') represents the incorrect assignment probability of the target token.
5. A two-stage few-shot automatic fact-checking method according to claim 4, wherein In step S4, select the top five sentences as the evidence sentences for claim verification.
6. The two-stage few-shot automatic fact-checking method according to claim 5, characterized in that The specific implementation method of step S5 includes the following steps: S5.
1. Select Flan-T5 as the base model of the pre-trained language model; S5.
2. Fine-tune the base model of the pre-trained language model using the LoRA fine-tuning method. The rank of the matrix is specified as 4 during training, and LoRA is applied to the q, k, v, o, and w layers of the SelfAttention, EncDecAttention, and DenseReluDense modules of the original model to obtain the pre-trained language model.
7. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of a two-stage few-shot automatic fact-checking method according to any one of claims 1-6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a two-stage few-shot automatic fact-checking method according to any one of claims 1-6.
Citation Information
Patent Citations
Automatic fact verification method fusing graph converter and common attention network
CN114048286A
False news detection method and system
CN116467530A