False news detection method based on multi-source retrieval and large language model
By combining text corpus and knowledge graph for multi-source search, and using large language models to generate abstracts and supplementary evidence, the shortcomings of existing fake news detection methods in obtaining and exploiting evidence are solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202510175770.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
Existing false news detection methods rely on a single source of knowledge for external evidence retrieval, and the inability to obtain more comprehensive evidence, and directly entering external evidence into large language models may lead to the model's over-reliance on external evidence and neglecting internal knowledge.
Multi-source search method is used to search external evidence in combination with text corpus and knowledge graphs, and abstracts and supplementary evidence are generated through large language models to improve the accuracy of fake news detection.
Obtain more comprehensive external evidence through multi-source search and effectively utilize internal knowledge of large language models by generating abstracts and supplementing evidence, thereby improving the accuracy of fake news detection.
Smart Images

Figure CN120124740A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information dissemination security, and more specifically, relates to a method for detecting false news based on multi-source retrieval and large language models. Background Art
[0002] The rapid progress of the Internet has greatly facilitated information acquisition. However, this trend has also accelerated the rapid spread of false news. These false news often have misleading, deceptive, and even malicious elements, which may mislead the public's perception, cause social chaos, and deeply affect public opinion, thereby triggering a series of negative consequences. The earliest false news detection methods mainly relied on time-consuming and laborious manual verification. Therefore, developing effective automated false news detection methods is crucial for suppressing the spread of false news and ensuring the authenticity of information.
[0003] Early automated false news detection research extracted features of news content, such as language features, writing styles, and sentiment patterns, through neural networks to train models to predict the authenticity of news. Although these methods have achieved certain results in automated false news detection, they mainly rely on the features learned by the models themselves for prediction, while ignoring the external evidence information related to the news. In contrast, false news detection methods based on fact-checking use relevant external evidence as additional input, which can not only improve the detection performance but also provide reliable support for the detection results. Therefore, fact-checking-based methods have become the main research direction in the current false news detection field. Although these methods have made significant progress after years of development, they still require a large amount of manually labeled data during the training process. In addition, the results predicted by the models are usually just a simple classification label, which is not conducive to people's understanding of the reasons for the models' predictions.
[0004] Recently, large language models (LLMs) have shown great potential in multiple tasks. These models are pre-trained on a large amount of data and world knowledge, giving them significant advantages in fake news detection. On the one hand, large language models do not require a large amount of domain-specific labeled data for training and have strong generalization capabilities. On the other hand, thanks to the excellent text generation ability of large language models, they can generate easy-to-understand explanations to help people understand the prediction results. Currently, fact-checking-based fake news detection methods that combine large language models usually adopt few-shot prompts learning to guide the model to decompose complex statements into simpler sub-statements or sub-questions, then search for external evidence in a specific corpus or on the Internet based on these sub-statements or sub-questions, and finally verify the sub-statements or sub-questions one by one according to this evidence and generate interpretable judgment reasons. Although these methods effectively utilize the powerful reasoning and text generation capabilities of large language models, there are still some deficiencies:
[0005] 1. Existing methods rely on a single knowledge source for external evidence retrieval, unable to obtain more comprehensive evidence, which limits the model's ability to verify statements from multiple perspectives.
[0006] 2. Existing methods usually directly input the retrieved external evidence into the large language model to verify the authenticity of the statement, but this method has certain defects. First, a large amount of redundant information is contained in the retrieved external evidence. If not properly processed, it will affect the accuracy of the model's judgment. Second, existing methods retrieve external evidence by decomposing the statement, but the decomposed sub-statements or sub-questions may deviate from the original statement, thus introducing more irrelevant information in the retrieval stage. In addition, directly inputting external evidence to LLMs may cause the model to overly rely on these input evidences during the judgment process and ignore its rich internal knowledge reserve. Summary of the Invention
[0007] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a fake news detection method based on multi-source retrieval and large language models, which combines text corpus retrieval and knowledge graph retrieval, and generates summaries and supplementary evidence with the help of large language models, thereby improving the accuracy of fake news detection.
[0008] To achieve the above-mentioned invention purpose, the fake news detection method based on multi-source retrieval and large language models of the present invention includes the following steps:
[0009] S1: Pre-obtain a text corpus as needed in advance, and retrieve the N pieces of corpus with the greatest relevance to the news C to be detected in it to form a relevant corpus set X = {x 1 ,x 2 ,…,xN}, x n represents the nth relevant corpus of the news C to be detected obtained from the text corpus, where n = 1, 2, …, N, and the value of N is set according to actual needs;
[0010] S2: Obtain a knowledge graph in advance according to actual needs, and retrieve the news C to be detected in the knowledge graph to generate a set of relevant corpora Y = {y 1 , y 2 , …, y M}, where y m represents the mth relevant corpus of the news C to be detected obtained from the knowledge graph, where m = 1, 2, …, M, and the value of M is set according to actual needs; The specific method of knowledge graph retrieval is as follows:
[0011] S2.1: Use a pre-trained large language model to extract entities from the news C to be detected, and denote the entity set E = {e 1 , e 2 , …, e K}, where e k represents the kth entity in the news C to be detected, where k = 1, 2, …, K, and K represents the number of entities in the news C to be detected;
[0012] S2.2: For each entity e k , retrieve its knowledge graph path set P k in the knowledge graph. The specific method is as follows:
[0013] S2.2.1: Initialize the entity set entity 1 = {e k};
[0014] S2.2.2: Let the retrieval round g = 1;
[0015] S2.2.3: For each entity g i = 1, 2, …, B in the entity set entity g , where B g represents the number of entities in the entity set entity g , retrieve the entity-relationship-target entity triples in the knowledge graph to obtain the retrieval result where represents the dth triple obtained from the entity , represents the target entity associated with the entity , represents the relationship between the entity and ; Represents an entity The number of triples retrieved from the knowledge graph;
[0016] S2.2.4: For each entity in the entity set entity g in the entity set of triples, use the large language model to select the Q triples with the greatest relevance to the news C to be detected as the preferred triples, and obtain the preferred retrieval result where represents an entity The q-th preferred triple obtained, q = 1, 2,..., Q, and the value of Q is set according to actual needs;
[0017] S2.2.5: Determine whether the round g < G, where G represents the preset maximum number of rounds. If so, go to step S2.2.6; otherwise, go to step S2.2.8;
[0018] S2.2.6: For the Q preferred triples of each entity in the entity set entity g in the entity set Construct the entity set entity by combining B g ×Q target entities ; g+1 ;
[0019] S2.2.7: Let the retrieval round g = g + 1, and return to step 2.2.3;
[0020] S2.2.8: Integrate the retrieval results of G rounds to generate all the knowledge graph paths corresponding to the entity e k to obtain its knowledge graph path set P k , and each knowledge graph path p k,h =(e k , r k,h,1 , t k,h,1 , r k,h,2 , t k,h,2 ,…, r k,h,Q , t k,h,G ), where t k,h,g represents the g-th hop target entity associated with the entity e k , r k,h,g represents the relationship between the entity t k,h,g-1 and t k,h,g , t k,h,0 =e k , h = 1, 2,..., H k , H k represents the number of knowledge graph paths generated by the entity e k ;
[0021] S2.3: Merge the knowledge graph path sets P of K entities e k to obtain the knowledge graph path set P of the news C to be detected k , and then use a large language model to screen out the M knowledge graph paths with the greatest relevance to the news C to be detected from the knowledge graph path set P C as the preferred knowledge graph paths, obtaining the preferred knowledge graph path set C where denotes the m-th preferred knowledge graph path;
[0022] S2.4: Use a large language model to generate natural language text for each preferred knowledge graph path in the preferred knowledge graph path set as the relevant corpus y m ;
[0023] S3: Merge the relevant corpus sets X = {x 1 , x 2 , …, x N} and the relevant corpus set Y = {y 1 , y 2 , …, y M} to obtain the relevant corpus set Z = X ∪ Y, and use a large language model to generate several summaries of the relevant corpus set Z, thus forming the summary set S;
[0024] S4: Use the news C to be detected as the claim, and the summary set S as the evidence, and use a large language model to generate several supplementary evidences, thus forming the supplementary evidence set L;
[0025] S5: Use the news C to be detected as the claim, and the summary set S and the supplementary evidence set L as the evidence, and use a large language model to reason about the authenticity of the news C to be detected, and generate a prediction label on whether the news C to be detected is false.
[0026] The false news detection method based on multi-source retrieval and large language model of the present invention first retrieves the relevant corpus set with the greatest relevance to the news to be detected in the text corpus, then generates the relevant corpus set of the news to be detected according to the knowledge graph retrieval, merges the two relevant corpus sets and uses a large language model to generate the summary set, then uses the news to be detected as the claim, the summary set as the evidence, and uses a large language model to generate several supplementary evidences, and finally uses the news to be detected as the claim, the summary set and the supplementary evidence set as the evidence, and uses a large language model to obtain the prediction label on whether the news to be detected is false.
[0027] The present invention has the following beneficial effects:
[0028] 1. Instead of decomposing the news to be detected, the present invention directly performs multi-source knowledge retrieval based on the original news to be detected, so as to retrieve more comprehensive external evidence related to the news to be detected;
[0029] 2. The present invention filters redundant information by generating summaries and effectively utilizes the rich world knowledge inside the large language model through knowledge conversion of the large language model, thereby improving the accuracy of fake news detection. Description of the Drawings
[0030] Figure 1 is a flowchart of the specific implementation of the fake news detection method based on multi-source retrieval and large language model of the present invention;
[0031] Figure 2 is a flowchart of knowledge graph retrieval in the present invention;
[0032] Figure 3 is a flowchart of retrieving entity paths in the present invention;
[0033] Figure 4 is a flowchart of fake news detection in this embodiment;
[0034] Figure 5 is a comparison chart of the fake news detection accuracy of each experiment in the ablation experiment of this embodiment. Specific Embodiment
[0035] The following describes the specific implementation of the present invention with reference to the drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0036] Embodiment
[0037] Figure 1 is a flowchart of the specific implementation of the fake news detection method based on multi-source retrieval and large language model of the present invention. As Figure 1 shown, the specific steps of the fake news detection method based on multi-source retrieval and large language model of the present invention include:
[0038] S101: Retrieve relevant text corpus:
[0039] According to actual needs, a text corpus is pre-obtained, and N pieces of corpus with the greatest relevance to the news C to be detected are retrieved from it to form a relevant corpus set X = {x 1 , x 2 , …, x N}, x nDenote the nth relevant corpus of the news C to be detected based on the text corpus, where n = 1, 2, …, N, and the value of N is set according to actual needs. In this embodiment, the news text corpus uses the Wikipedia corpus, and the text corpus retrieval method uses the BM25 retrieval algorithm.
[0040] S102: Knowledge graph retrieval:
[0041] In addition to text corpus retrieval, the present invention also introduces a knowledge graph to retrieve more comprehensive evidence, that is, a knowledge graph is pre-obtained according to actual needs, and the news C to be detected is retrieved in the knowledge graph to generate a set of relevant corpora Y = {y 1 , y 2 , …, y M}, where y m denotes the mth relevant corpus of the news C to be detected based on the knowledge graph, where m = 1, 2, …, M, and the value of M is set according to actual needs. Figure 2 is the flowchart of the knowledge graph retrieval in the present invention. As Figure 2 shown, the specific steps of the knowledge graph retrieval in the present invention include:
[0042] S201: Extract entities:
[0043] Use a pre-trained large language model to extract entities from the news C to be detected, and denote the entity set E = {e 1 , e 2 , …, e K}, where e k denotes the kth entity in the news C to be detected, where k = 1, 2, …, K, and K represents the number of entities in the news C to be detected.
[0044] S202: Retrieve entity paths:
[0045] For each entity e k , retrieve its knowledge graph path set P k in the knowledge graph. Figure 3 is the flowchart of retrieving entity paths in the present invention. As Figure 3 shown, the specific steps of retrieving entity paths in the present invention include:
[0046] S301: Initialize the entity set entity 1 = {e k}.
[0047] S302: Let the retrieval round g = 1.
[0048] S303: Retrieve triples:
[0049] For the entity set entity gEach entity in i = 1, 2, …, B g , B g represents the number of entities in the entity set entity g . Retrieve the entity-relationship-target entity triples in the knowledge graph to obtain the retrieval result where represents the entity the d-th triple obtained represents the target entity associated with the entity , represents the entity and the relationship between represents the entity the number of triples retrieved in the knowledge graph
[0050] S304: Retrieval result screening:
[0051] For each entity in the entity set entity g in of the triples, use the large language model to screen out the Q triples with the greatest relevance to the news C to be detected as the preferred triples, obtaining the preferred retrieval result where represents the entity the q-th preferred triple obtained, q = 1, 2, …, Q, and the value of Q is set according to actual needs
[0052] S305: Determine whether the round g < G, where G represents the preset maximum number of rounds. If so, go to step S306; otherwise, go to step S308
[0053] S306: Generate entity set:
[0054] For each entity in the entity set entity g in of the Q preferred triples Construct the B g × Q target entities to obtain the entity set entity g+1 .
[0055] S307: Let the retrieval round g = g + 1, and return to step S303
[0056] S308: Generate path:
[0057] Integrate the retrieval results of G rounds to generate the entity ek All the corresponding knowledge graph paths are obtained to get its knowledge graph path set P k , and for each knowledge graph path p k,h =(e k , r k,h,1 , t k,h,1 , r k,h,2 , t k,h,2 ,…, r k,h,Q , t k,h,G ), where t k,h,g represents the g-th hop target entity associated with the entity e k , r k,h,g represents the relationship between the entities t k,h,g-1 and t k,h,g , t k,h,0 =e k , h = 1, 2, …, H k , and H k represents the number of knowledge graph paths generated by the entity e k .
[0058] S203: Knowledge graph path screening:
[0059] The knowledge graph path sets P k of K entities e k are merged to obtain the knowledge graph path set P C of the news C to be detected. Then, a large language model is used to screen out the M knowledge graph paths with the highest relevance to the news C to be detected from the knowledge graph path set P C as the preferred knowledge graph paths, obtaining the preferred knowledge graph path set where represents the m-th preferred knowledge graph path.
[0060] S204: Generate relevant corpus:
[0061] A large language model is used to generate natural language text as the relevant corpus y for each preferred knowledge graph path in the preferred knowledge graph path set m .
[0062] In this embodiment, the Wikidata knowledge graph is used, and the GPT-3.5-Turbo model is used as the large language model. The number of preferred triples Q = 2, and the number of retrieval rounds G = 2.
[0063] S103: Generate summary:
[0064] The relevant corpus set X = {x 1 , x 2 ,…, xN} and the related corpus set Y = {y 1 , y 2 , …, y M} are merged to obtain the related corpus set Z = X ∪ Y. A large language model is used to generate several abstracts of the related corpus set Z, thus forming an abstract set S. Redundant information in the related corpus set can be filtered by generating abstracts.
[0065] S104: Generate supplementary evidence:
[0066] Taking the news C to be detected as a statement and the abstract set S as evidence, a large language model is used to generate several pieces of supplementary evidence, thus forming a supplementary evidence set L. Through this step, the rich world knowledge inside the large language model can be replaced with additional supplementary evidence to improve the accuracy of fake news detection. In this embodiment, the number of supplementary evidences is set to 3.
[0067] S105: News detection:
[0068] Taking the news C to be detected as a statement and the abstract set S and the supplementary evidence set L as evidence, a large language model is used to reason about the authenticity of the news C to be detected and generate a prediction label on whether the news C to be detected is fake.
[0069] Next, a specific example is used to illustrate the specific process of the method of the present invention. In this embodiment, the text corpus uses the Wikipedia corpus, the knowledge graph uses the Wikipedia knowledge graph, and the large language model uses the GPT-3.5-Turbo model. The content of the news C to be detected is set to “The Kentucky Department of Corrections is headquartered along the Kentucky River.” Figure 4 This is the flowchart of fake news detection in this embodiment. As Figure 4 shown, first, several relevant corpora are searched in the text corpus, and then knowledge graph retrieval is performed. The specific process is as follows:
[0070] The entity set of the news C to be detected is extracted, including {Kentucky Department of Corrections, Kentucky River}. The prompt for entity extraction is as follows:
[0071] [Guidance]
[0072] Given a claim, if I want to verify the truth or falseness of the claim, help me extract the entities of the claim to be more suitable for knowledge graph to search for evidence. The entities should be short and no more than 5. Only entities such as proper names, places, and person need to be extracted, ignoring entities such as time, data, numbers, country, and verbs. If there is no suitable entity just answer None.
[0073] [Input]
[0074] Claim: {claim}
[0075] [Output Format]
[0076] Entities: ["Entity 1", "Entity 2"...] Remember to follow the format for output.
[0077] Retrieve the knowledge graph paths of each entity in the Wikipedia knowledge graph. The following are the tips for filtering the most relevant preferred triples (the number of preferred triples in this embodiment is 3) in the knowledge graph path retrieval:
[0078] [Guidance]
[0079] Based on the claim, extract the most relevant results from the following search results and only return their indices, with no more than three indices:
[0080] [Input]
[0081] Claim: {claim}
[0082] Search results: {combined strings}
[0083] Idx: ["idx1",...]
[0084] [Output Format]
[0085] Idx: ["idx1"] Remember to follow the format for output.
[0086] Tips for screening the knowledge graph path with the greatest relevance are as follows:
[0087] [Guidance]
[0088] Based on the results retrieved from the following knowledge graph, choose some of the most relevant path to claim and only return their indices:
[0089] [Input]
[0090] Claim: {claim}
[0091] Results: {result chain}
[0092] [Output Format]
[0093] Idx: ["idx1",...] Remember to follow the format for output. For example: Idx: ["idx1", "idx2", "idx3"]
[0094] Combine the relevant corpus set X obtained from the text corpus with the relevant corpus set Y obtained from the knowledge graph to obtain the relevant corpus set Z, and then use the large language model to generate the abstract.
[0095] [Guidance]
[0096] As a paragraph - summarizing assistant, you are required to complete the task according to the following rules: 1. Extract information related to the claim from the provided paragraphs and summarize it; the summary should not exceed 300 words. 2. Summarize all the information in the paragraphs that is relevant to the claim, but do not generate a summary based directly on the claim, as the claim may be incorrect. 3. Summary should be accurate and comprehensive. 4. Do not summarize irrelevant information. 5. Do not generate information that is not relevant to summary.
[0097] [Input]
[0098] Claim: {claim}
[0099] Paragraph1: {paragraph1}
[0100] Paragraph2: {paragraph2}
[0101] Then, take the news C to be detected as the claim and the summary set S as the evidence, and use a large - language model to generate a supplementary evidence set L. The prompts used are as follows:
[0102] [Guidance]
[0103] As an Information Retrieval Assistant, you are required to complete the task according to the following rules: 1. Based on the given information, retrieve 3 additional pieces of information that can help determine the correctness of the claim. 2. Each piece of additional information should not exceed 100 words. 3. Do not generate additional information directly based on the claim, as the claim may be incorrect. 4. If there is an error in the given information, please provide the correct information as additional information.
[0104] [Input]
[0105] Claim: {claim}
[0106] Summary: {summary}
[0107] [Output Format]
[0108] Additional information:
[0109] Use the news C to be detected as the claim, and the summary set S and the supplementary evidence set L as evidence. Use a large language model to obtain the prediction label of whether the news C to be detected is false. The prompts used are as follows:
[0110] [Guidance]
[0111] You are a CLAIM VERIFICATION ASSISTANT and you need to determine if the claim is correct based on the given evidence: 1. Based on the given information, retrieve 3 additional pieces of information that can help determine the correctness of the claim. 2. Each piece of additional information should not exceed 100 words. 3. Do not generate additional information directly based on the claim, as the claim may be incorrect. 4. If there is an error in the given information, please provide the correct information as additional information.
[0112] [Input]
[0113] Claim: {claim}
[0114] Evidence 1: {summary}
[0115] Evidence 2: {llm evidence}
[0116] [Output Format]
[0117] Based on the above evidences, is it true that claim?Please provide the answer (true / false) and the reason.
[0118] answer:
[0119] reason:
[0120] According to the prompt, in this embodiment, a reason text is also generated. This text is interpretable information, so as to explain the reason for the judgment of false news.
[0121] To better illustrate the technical effects of the present invention, specific examples are used to conduct experimental verification on the present invention. In this embodiment, 4 existing false news detection methods are selected for comparison, including:
[0122] The ProgramFC algorithm, see the literature "Pan L, Wu X, Lu X, et al. Fact-Checking Complex Claims with Program-Guided Reasoning[C] / / Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics(Volume 1: Long Papers). 2023: 6981-7004.";
[0123] The FOLK algorithm, see the literature "Wang H, Shu K. Explainable Claim Verification via Knowledge-Grounded Reasoning with Large Language Models[C] / / The 2023 Conference on Empirical Methods in Natural Language Processing.";
[0124] The HiSS algorithm, see the literature "Zhang X, Gao W. Towards LLM-based Fact Verification on News Claims with a Hierarchical Step-by-Step Prompting Method[C] / / Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics(Volume 1: Long Papers). 2023: 996-1011.";
[0125] For the CoK algorithm, see Document 4: "Li X, Zhao R, Chia Y K, et al. Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous Sources [C] / / The Twelfth International Conference on Learning Representations."
[0126] Next, experiments were conducted on the present invention (denoted as IMRRF) and the comparative methods in terms of the detection accuracy (Accuracy), F1 score (F1 Score), precision (Precision), recall (Recall), etc. Table 1 shows the comparison results of the present invention and the comparative methods on different datasets and different evaluation metrics.
[0127]
[0128] Table 1
[0129] It can be seen from Table 1 that:
[0130] (1) Compared with HiSS that directly uses Few-Shot Prompts Learning to decompose complex statements into sub-statements, FOLK more effectively decomposes sub-statements by extracting predicates from the statements. ProgramFC better utilizes the reasoning ability of large language models by converting statements into logical reasoning programs. Both FOLK and ProgramFC are superior to HiSS in detection performance, indicating that the quality of statement decomposition directly affects the detection effect of the model. In contrast, the present invention skips the step of statement decomposition and directly retrieves external evidence based on the statements, thus avoiding the influence of decomposition quality on the detection results and achieving excellent detection performance.
[0131] (2) Compared with HiSS and FOLK that only rely on web page retrieval, CoK combines web page retrieval and knowledge graph retrieval to obtain external evidence. Its overall performance exceeds that of HiSS and FOLK, demonstrating the advantages of multi-source knowledge retrieval and indicating that knowledge graph queries can retrieve more reliable evidence.
[0132] (3) On the FEVEROUS dataset, the performance of the present invention exceeds that of ProgramFC, with the accuracy rate increasing by 8.38% and the F1 score increasing by 9.67%. Similarly, on the HOVER dataset, the proposed method performs better. Especially when dealing with 3-hop and 4-hop claims, compared to ProgramFC, the accuracy rate increases by an average of 5.9% and the F1 score increases by 7.5%. These results indicate that IMRRF integrates evidence from specific corpora and knowledge graphs during the external evidence retrieval process, while effectively utilizing the internal knowledge of large language models as supplementary evidence. The present invention provides more comprehensive and accurate evidence, enabling the model to achieve better performance.
[0133] (4) When dealing with the HOVER-2hop dataset, the detection performances of methods such as HiSS, FOLK, and CoK are similar. Even the difference between the best baseline method ProgramFC and the present invention is not significant. This is because when dealing with simple claims like 2-hop, the evidence retrieved by different methods does not have significant differences, resulting in small differences in detection performance. However, when dealing with more complex claims, the present invention provides more comprehensive and accurate evidence, thus having a greater advantage in detection performance.
[0134] (5) In the FEVEROUS dataset, the present invention shows a balanced performance in terms of precision, recall, and F1 score in the refutation and support categories, and the F1 score is significantly higher than other methods. This indicates that the present invention can find more evidence while maintaining a high accuracy rate. Similarly, in the HOVER dataset, the present invention maintains a performance advantage in the refutation category. In the support category, especially for more complex 3-hop and 4-hop claims, the present invention continues to perform well in terms of the F1 score, especially with stable recall, highlighting its superior adaptability in dealing with complex claims.
[0135] To verify the effectiveness of the present invention in knowledge graph retrieval, generating summaries, and generating supplementary evidence, ablation experiments are designed in this embodiment. Figure 5 It is a comparison chart of the false news detection accuracy rates of each experiment in the ablation experiment of this embodiment. As Figure 5 shown, adding knowledge graph retrieval on the basis of text corpus retrieval can increase the accuracy rates of the FEVEROUS and HOVER-2hop datasets by 0.4% and 1.34% respectively. However, on the HOVER-3hop and HOVER-4hop datasets, the accuracy rates decrease by 2.4% and 0.77% respectively, probably because it is difficult for the model to filter out irrelevant information when integrating evidence, thus affecting its verification ability.
[0136] Then, experiments were conducted on the combination of the text corpus and the knowledge graph retrieval results. Compared with generating summaries only based on the evidence retrieved from the text corpus, adding the knowledge graph retrieval and then performing the summary generation step increased the detection accuracy by 1.05%, 3.29%, 0.39%, and 0.1% respectively on each dataset, indicating that the knowledge graph can effectively retrieve more relevant evidence. At the same time, the evidence obtained by generating summaries can effectively filter out irrelevant information.
[0137] Next, the evidence retrieved from the text corpus and the knowledge graph was used to guide the large language model to perform knowledge transformation to generate supplementary evidence, and experiments were conducted with this generated supplementary evidence as an additional input. In the FEVEROUS dataset, using the large language model to generate supplementary evidence increased the accuracy by 3% compared to not using this evidence. In the HOVER dataset, the accuracy increased by 3.28%, 6.54%, and 3.08% respectively. This shows that the large language model can generate valuable supplementary evidence. Similarly, the results obtained by using the evidence retrieved from the text corpus and the knowledge graph to generate summaries were used to guide the large language model to generate supplementary evidence and use it as an additional input. In the FEVEROUS dataset, this method increased the accuracy by 4.93%; in the HOVER dataset, the accuracy increased by 0.53%, 4.09%, and 2.53% respectively. These results further verified the effectiveness of the large language model in generating supplementary evidence through knowledge transformation.
[0138] Although the above describes the illustrative specific embodiments of the present invention for the convenience of those skilled in the art to understand the present invention, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A fake news detection method based on multi-source retrieval and large language model, characterized in that: The following steps are involved: S1: Pre-acquire a text corpus according to actual needs, and retrieve the N pieces of text that are most relevant to the news C to be detected to form a relevant text set X = {x1, x2, ..., x N }, x n represents the nth relevant corpus of the news C to be detected obtained based on the text corpus, n = 1, 2, ..., N, and the value of N is set according to actual needs; S2: Pre-acquire the knowledge graph according to actual needs, search the news C to be detected in the knowledge graph, and generate the relevant corpus set Y = {y1, y2, ..., y M },y m represents the mth relevant corpus of the news C to be detected based on the knowledge graph, m = 1, 2, ..., M, and the value of M is set according to actual needs; the specific method of knowledge graph retrieval is: S2.1: Use the pre-trained large language model to extract entities from the news to be detected C, and record the entity set E = {e1, e2, …, e K },e k represents the kth entity in the news to be detected C, k = 1, 2, ..., K, K represents the number of entities in the news to be detected C; S2.2: For each entity e k , retrieve in the knowledge graph to obtain its knowledge graph path set P k , the specific method is: S2.2.1: Initialize entity set entity1 = {e k }; S2.2.2: Let the search round g = 1; S2.2.3: For entity collection entity g Each entity i=1,2,…,B g , B g Represents entity collection entity g The number of entities in the knowledge graph is retrieved to obtain the entity-relationship-target entity triples and obtain the search results in Representing Entities The dth triplet obtained is, Representation and Entity The associated target entity, Representing Entities and The relationship between Representing Entities The number of triples retrieved in the knowledge graph; S2.2.4: For entity collection entity g Each entity of The large language model is used to select Q triples with the greatest relevance to the news C to be detected as the preferred triples to obtain the preferred search results. in Representing Entities The qth optimal triplet is obtained, q = 1, 2, ..., Q, and the value of Q is set according to actual needs; S2.2.5: Determine whether the round g < G, where G represents the preset maximum round. If so, proceed to step S2.2.6; otherwise, proceed to step S2.2.8; S2.2.6: For entity collection entity g Each entity Q preferred triplets B g ×Q target entities Construct the entity set entity g+1 ; S2.2.7: Set the search round g = g + 1, and return to step 2.2.3; S2.2.8: Integrate the search results of round G to generate entity e k All the corresponding knowledge graph paths, thus obtaining its knowledge graph path set P k , each knowledge graph path p k,h =(e k ,r k,h,1 ,t k,h,1 ,r k,h,2 ,t k,h,2 ,…,r k,h,Q ,t k,h,G ), where t k,h,g Represents entity e k The associated g-th hop target entity, r k,h,g Represents entity t k,h,g-1 and t k,h,g The relationship between k,h,0 =e k ,h=1,2,…,H k , H k Represents entity e k The number of knowledge graph paths generated; S2.3: K entities e k The knowledge graph path set P k Merge to get the knowledge graph path set P of the news C to be detected C , and then use the large language model to extract the path set P from the knowledge graph C The M knowledge graph paths with the greatest relevance to the news C to be detected are selected as the preferred knowledge graph paths, and the preferred knowledge graph path set is obtained. in represents the mth preferred knowledge graph path; S2.4: Use a large language model to select the set of knowledge graph paths Each preferred knowledge graph path in Generate natural language text as relevant corpus y m ; S3: The relevant corpus set X = {x1, x2, ..., x N } and the related corpus set Y = {y1,y2,…,y M Merge to obtain a related corpus set Z = X∪Y, and use a large language model to generate several summaries of the related corpus set Z, thereby forming a summary set S; S4: Take the news to be detected C as a statement, the summary set S as evidence, and use the large language model to generate several supplementary evidences to form a supplementary evidence set L; S5: Take the news C to be tested as a statement, the summary set S and the supplementary evidence set L as evidence, use the large language model to infer the authenticity of the news C to be tested, and generate a prediction label whether the news C to be tested is false.
2. The false news detection method according to claim 1, characterized in that: The text corpus in step S1 adopts the Wikipedia corpus.
3. The false news detection method according to claim 1, characterized in that: The text corpus retrieval method in step S1 adopts the BM25 retrieval algorithm.
4. The false news detection method according to claim 1, characterized in that: The knowledge graph in step S2 adopts the Wikidata knowledge graph.
5. The false news detection method according to claim 1, characterized in that: In step S2, the large language model adopts the GPT-3.5-Turbo model.
Citation Information
Cited By
Manuscript generation service auxiliary system based on AI intelligent large model
CN120781807A