A rumor detection system and method based on large language models

The core contradictions in the rumor text were extracted through a large language model and combined with knowledge base and Internet search, the problem that existing rumor detection methods cannot effectively utilize external facts is solved, and efficient and interpretable rumor detection is achieved.

CN120181073BActive Publication Date: 2025-07-22CHANGSHA ZHIWEI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510648140.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-22
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing rumor detection methods cannot effectively utilize external objective facts and are poorly interpretable, making it difficult to distinguish the core contradictions in rumor information.

Method used

A rumor detection system based on a large language model is adopted, including a core contradiction extraction module, a fact verification search module and a rumor judgment decision-making module. The entity relationship and attribute core contradictions are extracted through a large language model, combined with the knowledge base and the Internet to search related credible judgment basis, the impact of each judgment basis is analyzed to form an interpretable detection result.

Benefits of technology

It improves the accuracy and interpretability of rumor detection, reduces interference from irrelevant information, enriches the source of judgment basis, and improves information utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181073B_ABST
    Figure CN120181073B_ABST
Patent Text Reader

Abstract

The present invention discloses a rumor detection system based on a large language model, which includes a core contradiction extraction module, a fact-checking retrieval module, and a rumor judgment and decision-making module. By extracting the core contradiction through the large language model, the present invention can reduce the probability of being interfered by irrelevant information, facilitate subsequent retrieval, improve the information utilization efficiency, and enhance the rumor detection accuracy. At the same time, through the fact-checking retrieval module, relevant judgment bases are retrieved from the knowledge base and the Internet, enriching the source of judgment bases, while avoiding the influence of hallucinations of the large language model on the detection results, and further improving the reliability of the judgment bases through the filtering of the information source database. Finally, the rumor judgment and decision-making module enables the large language model to analyze the influence of each judgment basis summary on the core contradiction, form a detection result, simulates the process of human judgment, makes the detection result based on the judgment basis, and has interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information data processing, and particularly relates to a rumor detection system and method based on a large language model. Background Art

[0002] There is a vast amount of information on the Internet, and at the same time, the information evolves rapidly, making it inefficient for manual judgment and rumor refutation. Therefore, automated rumor detection methods are needed to identify rumors from the vast amount of information on the Internet and the network environment.

[0003] Most of the early automated rumor identification technologies were based on text features and user metadata, including less data information, so the detection accuracy was insufficient. Subsequently, Bian et al. proposed a detection method based on the rumor propagation structure, using a graph convolutional network to traverse from two directions of propagation, thereby capturing the features in the process of rumor text propagation and improving the rumor detection efficiency. Since Internet information is not limited to text data, but also includes various modalities such as pictures, audio, and video, Chen et al. separately abstracted multiple modalities such as text and images and used the goal of reinforcement learning for optimization, improving the rumor detection ability for multi-modal information.

[0004] These rumor detection methods objectively improve the rumor detection efficiency, but their detection methods are all based on the information given by rumor fabricators, and these information often contain other facts used to confuse and interfere with the detection. Existing detection methods are difficult to effectively utilize external objective facts to distinguish the core contradictions from rumor information. At the same time, the existing rumor detection methods have poor interpretability of the detection results, reducing the persuasiveness of the detection results. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the defects that the existing rumor detection methods cannot effectively utilize external objective facts and have poor interpretability, so as to provide a rumor detection system and method based on a large language model.

[0006] The present invention provides a rumor detection system based on a large language model, including:

[0007] A core contradiction extraction module, configured to: input a first prompt word and a text to be detected into the large language model, and extract the core contradiction of entity relationships in the text to be detected, where the core contradiction of entity relationships is the contradiction between the mutual relationships of multiple entities in the text to be detected output by the large language model and the facts; input a second prompt word and the text to be detected into the large language model, and extract the core contradiction of entity attributes in the text to be detected, where the core contradiction of entity attributes is the contradiction between the attributes of entities in the text to be detected output by the large language model and the facts;

[0008] The fact-checking retrieval module includes a knowledge base retrieval module and an Internet retrieval module; the knowledge base retrieval module is used to: retrieve multiple pieces of knowledge base data most relevant to the core contradiction in the knowledge base through the RAG algorithm to obtain relevant knowledge base data; the Internet retrieval module is used to: retrieve the core contradiction from the Internet to obtain a retrieval result, filter the retrieval result through a source database to obtain a credible retrieval result; input the third prompt word, the core contradiction, and the credible retrieval result into a large language model to obtain relevant credible retrieval results; the core contradiction includes the entity relationship core contradiction and the entity attribute core contradiction;

[0009] The rumor judgment and decision-making module is used to: input the fourth prompt word, the core contradiction, and the judgment basis into a large language model, summarize each judgment basis, and extract the content that mutually corroborates or contradicts the core contradiction therein to form a judgment basis summary; the judgment basis includes the relevant knowledge base data and the relevant credible retrieval results; input the fifth prompt word, the text to be detected, the core contradiction, and the judgment basis summary into a large language model, and analyze the influence of each judgment basis summary on the core contradiction to form a detection result.

[0010] Further, the core contradiction extraction module is also used to: add its time mark when extracting the entity relationship core contradiction in the text to be detected, and add its time mark when extracting the entity attribute core contradiction in the text to be detected;

[0011] The fact-checking retrieval module is also used to: when retrieving and obtaining relevant knowledge base data, filter out the knowledge base data whose time is completely irrelevant to the corresponding time mark; when retrieving the core contradiction from the Internet to obtain a retrieval result, filter out the retrieval result whose time is completely irrelevant to the corresponding time mark.

[0012] Further, the knowledge base retrieval module is also used to: crawl the data of the rumor refutation platform and establish a knowledge base.

[0013] Further, the knowledge base retrieval module is also used to:

[0014] Vectorize each piece of knowledge base data in the knowledge base to form vectorized knowledge base data; vectorize the entity relationship core contradiction and the entity attribute core contradiction to form a vectorized entity relationship core contradiction and a vectorized entity attribute core contradiction;

[0015] Calculate the similarity score between each vectorized entity relationship core contradiction and each vectorized knowledge base data:

[0016] ;

[0017] where, represents the vectorized entity relationship core contradiction, Indicates the th vectorized knowledge base data;

[0018] Calculate the similarity score between the core contradiction of each vectorized entity attribute and each vectorized knowledge base data:

[0019] ;

[0020] Among them, represents the core contradiction of the vectorized entity attribute, indicates the th vectorized knowledge base data;

[0021] Take the larger value between the similarity score of the core contradiction of the vectorized entity relationship and the vectorized knowledge base data and the similarity score of the core contradiction of the vectorized entity attribute and the vectorized knowledge base data as the final similarity score of the corresponding knowledge base data:

[0022] ;

[0023] Select multiple knowledge base data with the highest final similarity scores as the relevant knowledge base data.

[0024] Furthermore, the Internet retrieval module is also used for: requiring the large language model to classify the credible retrieval results into information with no inspection value, information with inspection value, and completely irrelevant information in the third prompt word; the information with no inspection value is an entity related to the core contradiction but has no direct relationship with the core contradiction, or the provided information is useless, untrustworthy, or irrelevant; the information with inspection value is related to the core contradiction and provides useful information or clues for further investigation; the completely irrelevant information is an entity that does not involve the core contradiction mentioned at all, or is completely irrelevant to the core contradiction, or is meaningless or misleading information.

[0025] Furthermore, the Internet retrieval module is also used for: establishing a source database and classifying the sources into high-credibility sources and low-credibility sources; the high-credibility sources include media and commercial websites; the low-credibility sources include individuals, self-media, and forums.

[0026] Further, the rumor judgment and decision-making module is further configured to: in the fifth prompt word: require the large language model to analyze whether the information in the judgment basis summary is consistent with the information in the text to be detected, and judge whether to affirm or deny the core contradiction; require the large language model to judge the reliability of the corresponding judgment basis based on the information source of each judgment basis summary; require the large language model to judge whether the text to be detected is a rumor based on the reliability of the judgment basis summary and the impact on the core contradiction; require the large language model to output the analysis and judgment process of each judgment basis summary on the core contradiction of the text to be detected, and point out at least one of the most important judgment basis summaries in the analysis and judgment process.

[0027] A rumor detection method based on a large language model using the above system includes:

[0028] Input the first prompt word and the text to be detected into the large language model to extract the core contradiction of the entity relationship in the text to be detected, where the core contradiction of the entity relationship is the contradiction between the mutual relationship of multiple entities in the text to be detected output by the large language model and the fact; input the second prompt word and the text to be detected into the large language model to extract the core contradiction of the entity attribute in the text to be detected, where the core contradiction of the entity attribute is the contradiction between the attribute of the entity in the text to be detected output by the large language model and the fact;

[0029] Retrieve multiple pieces of knowledge base data most relevant to the core contradiction in the knowledge base through the RAG algorithm to obtain relevant knowledge base data; retrieve the core contradiction from the Internet to obtain a retrieval result, and filter the retrieval result through the information source database to obtain a credible retrieval result; input the third prompt word, the core contradiction, and the credible retrieval result into the large language model to obtain a relevant credible retrieval result; the core contradiction includes the core contradiction of the entity relationship and the core contradiction of the entity attribute;

[0030] Input the fourth prompt word, the core contradiction, and the judgment basis into the large language model, summarize each judgment basis, and extract the content that corroborates or contradicts the core contradiction to form a judgment basis summary; the judgment basis includes the relevant knowledge base data and the relevant credible retrieval result; input the fifth prompt word, the text to be detected, the core contradiction, and the judgment basis summary into the large language model to analyze the impact of each judgment basis summary on the core contradiction to form a detection result.

[0031] A computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the above method.

[0032] A computer device, the computer device includes a memory, a processor, and a program stored and executable on the memory. When the program is executed by the processor, the above steps are implemented.

[0033] Beneficial effects: The present invention discloses a rumor detection system based on a large language model, including a core contradiction extraction module, a fact-checking retrieval module, and a rumor judgment and decision-making module. The core contradiction extraction module is used to extract the core contradiction of entity relationships and the core contradiction of entity attributes from the text to be detected through the large language model. The fact-checking retrieval module includes a knowledge base retrieval module and an Internet retrieval module, which are used to retrieve relevant credible judgment bases from the knowledge base and the Internet based on the core contradiction of entity relationships and the core contradiction of entity attributes. The rumor judgment and decision-making module is used to analyze and judge the impact of the judgment basis on the authenticity of the text to be detected through the large language model based on the core contradiction and the judgment basis, and finally output an analysis and demonstration and a detection result based on the judgment basis. By extracting the core contradiction through the large language model, the present invention can reduce the probability of being interfered by irrelevant information, facilitate subsequent retrieval, improve the information utilization efficiency, and enhance the rumor detection accuracy; at the same time, by retrieving relevant judgment bases in the knowledge base and the Internet through the fact-checking retrieval module, the source of the judgment basis is enriched, and at the same time, the hallucination of the large language model is avoided from affecting the detection result, and the reliability of the judgment basis is further improved through the filtering of the information source database; finally, the rumor judgment and decision-making module enables the large language model to analyze the impact of each judgment basis summary on the core contradiction to form a detection result, simulating the process of human judgment, making the detection result based on the judgment basis and having interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a schematic block diagram of the method flow of the present invention;

[0036] Figure 2 It is a schematic diagram of the method step flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To make the above objects, features, and advantages of the present application more apparent and understandable, the following describes the specific embodiments of the present application in detail with reference to the accompanying drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0038] In the description of the present application, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

[0039] In the present application, unless otherwise clearly specified and limited, terms such as "installed", "connected", "connected to", "fixed" and the like shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0040] Embodiment 1:

[0041] Referring to Figure 1 As shown, this embodiment provides a rumor detection system based on a large language model, including:

[0042] A core contradiction extraction module, configured to: input a first prompt word and a text to be detected into the large language model, and extract the core contradiction of entity relationships in the text to be detected, where the core contradiction of entity relationships is the contradiction between the mutual relationships of multiple entities in the text to be detected output by the large language model and the facts; input a second prompt word and the text to be detected into the large language model, and extract the core contradiction of entity attributes in the text to be detected, where the core contradiction of entity attributes is the contradiction between the attributes of the entity in the text to be detected output by the large language model and the facts;

[0043] Specifically, for the core contradiction of entity relationships, that is, the contradiction between the mutual relationships of multiple entities in the text to be detected and the facts, this embodiment designs the prompt words of the large language model, analyzes the mutual relationships between different entities in the text to be detected, and then judges whether there is a possibility of exaggeration, distortion, or fabrication in these interaction relationships, and analyzes whether there is a possibility of not conforming to the facts. In this embodiment, the first prompt word is:

[0044] {The following text may be a rumor. Please identify two or more entities involved and their relationships. Analyze whether the relationships between these entities conform to reality, and whether there are suspicions of exaggeration, distortion, or fabrication. Pay special attention to whether there are parts that do not conform to common sense or known facts. Ensure that all entities related to contradictions are involved in the analysis, and clearly indicate the time when the contradictions occur. Extract and summarize these parts to form a simple sentence that can describe the contradiction without summarizing whether it is reasonable:};

[0045] Input the above first prompt word and the text to be detected into the large language model, and the short sentence output by the large language model is the core contradiction of the entity relationship ;

[0046] For the core contradiction of entity attributes, that is, the contradiction between the attributes of the entities in the text to be detected and the facts, this embodiment designs prompt words for the large language model to analyze and clarify the key core entities in the text to be detected, and analyze whether there are situations of exaggeration, distortion, or fabrication, and find the key points and key attributes that may be rumors. In this embodiment, the second prompt word is:

[0047] {The following text may be a rumor. Please analyze the key entities mentioned in the text and their attribute descriptions, and judge whether these attributes conform to the actual situation and whether there are suspicions of exaggeration, distortion, or fabrication. Extract and summarize from these suspicious parts to form a short sentence, ensuring to describe the entity and the corresponding state of the entity, and clearly indicate the time when the contradiction occurs without summarizing whether it is reasonable:};

[0048] Input the above first prompt word and the text to be detected into the large language model, and the short sentence output by the large language model is the core contradiction of entity attributes .

[0049] The fact-checking retrieval module includes a knowledge base retrieval module and an Internet retrieval module; the knowledge base retrieval module is used to: retrieve multiple pieces of knowledge base data most relevant to the core contradiction in the knowledge base through the RAG algorithm to obtain relevant knowledge base data , and select k pieces of knowledge base data in this embodiment; the Internet retrieval module is used to: retrieve the core contradiction from the Internet to obtain a retrieval result, filter the retrieval result through the source database to obtain a credible retrieval result; input the third prompt word, the core contradiction, and the credible retrieval result into the large language model to obtain relevant credible retrieval results; the core contradiction includes the core contradiction of entity relationship and the core contradiction of entity attributes;

[0050] In this embodiment, the knowledge base retrieval module is further used to: crawl the data of the rumor-refuting platform, which is the China Internet Joint Rumor Refuting Platform (www.piyao.org.cn) in this embodiment, and establish a knowledge base;

[0051] The knowledge base retrieval module is further configured to:

[0052] Vectorize each piece of knowledge base data in the knowledge base to form vectorized knowledge base data ; Vectorize the core contradiction of entity relationship and the core contradiction of entity attribute to form vectorized core contradiction of entity relationship and vectorized core contradiction of entity attribute;

[0053] Calculate the similarity score between each vectorized core contradiction of entity relationship and each vectorized knowledge base data:

[0054] ;

[0055] Wherein, represents the vectorized core contradiction of entity relationship, represents the th piece of vectorized knowledge base data;

[0056] Calculate the similarity score between each vectorized core contradiction of entity attribute and each vectorized knowledge base data:

[0057] ;

[0058] Wherein, represents the vectorized core contradiction of entity attribute, represents the th piece of vectorized knowledge base data;

[0059] Take the larger value between the similarity score of the vectorized core contradiction of entity relationship and the vectorized knowledge base data and the similarity score of the vectorized core contradiction of entity attribute and the vectorized knowledge base data as the final similarity score of the corresponding knowledge base data:

[0060] ;

[0061] Select multiple pieces of knowledge base data with the highest final similarity scores as relevant knowledge base data; In this embodiment, select the k pieces of knowledge base data with the highest final similarity scores as relevant knowledge base data: ;

[0062] Specifically, the Internet retrieval module is configured to: call the search engine API interface to retrieve the core key contradiction from the Internet and , obtain web pages or documents related to the core contradiction, that is, retrieval results; read the web pages or documents given by the search engine one by one, filter out those from low-trust information sources according to the information source database, and retain other search results from high-trust information sources, that is, credible retrieval results;

[0063] Input the third prompt, the core contradiction, and the credible retrieval results into the large language model, so as to classify each credible retrieval result into three types: no inspection value, having inspection value, and completely irrelevant; no inspection value means that the content searched may involve the entity mentioned in the key contradiction, but has little relation to the key contradiction; having inspection value means that the information is indeed related to the key contradiction, provides effective information, and is worthy of further inspection; completely irrelevant means that the retrieval result does not involve the entity mentioned in the key contradiction at all, or has nothing to do with the key contradiction; in this embodiment, the third prompt is:

[0064] {The following is about a key contradiction <c>For the credible retrieval results, judge the relevance of each piece of information and classify it into three categories:

[0065] 1. "Of no inspection value": This piece of information may involve the entity mentioned in the key contradiction, but has no direct relation to the key contradiction, or the information provided is useless, untrustworthy or irrelevant.

[0066] 2. "Of inspection value": This piece of information is related to the key contradiction and provides useful information or clues for further investigation.

[0067] 3. "Completely irrelevant": This piece of information does not involve the entity mentioned in the key contradiction at all, or has no relation to the key contradiction, and is even meaningless or misleading information.

[0068] Please classify each piece of information according to the following search results:

[0069] Credible retrieval result 1: <Content of the credible retrieval result>

[0070] Credible retrieval result 2: <Content of the credible retrieval result>

[0071] ……

[0072] Core contradiction c: <Core contradiction>

[0073] Return the classification (of no inspection value, of inspection value, completely irrelevant) for each search result.};

[0074] Among them, the credible retrieval results of inspection value output by the large language model are the relevant credible retrieval results; reserve the relevant credible retrieval results for the core contradiction of entity relationship and the core contradiction of entity attribute respectively, that is, there are ;

[0075] Merge the relevant knowledge base data and the relevant credible retrieval results to form the basis for judgment, that is, there is .

[0076] The rumor judgment decision module is used to: input the fourth prompt word, the core contradiction, and the basis for judgment into the large language model, summarize each basis for judgment, and extract the content that mutually corroborates or contradicts the core contradiction therein to form a summary of the basis for judgment , in this embodiment, the summary of the basis for judgment marks the specific information source; the basis for judgment includes the relevant knowledge base data and the relevant credible retrieval results; input the fifth prompt word, the text to be detected, the core contradiction, and the summary of the basis for judgment into the large language model, and analyze the influence of each summary of the basis for judgment on the core contradiction to form a detection result;

[0077] In this embodiment, the fourth prompt word is:

[0078] The following are multiple data entries from the rumor refutation knowledge base and the Internet, as well as the key contradictions and . Please summarize each piece of data and extract the parts that may corroborate or contradict the key contradictions. The generated summary should include the following content:

[0079] 1. If the data is related to any of the key contradictions and , analyze whether the data provides evidence to support or refute the contradictions. Clearly mark whether the data corroborates or contradicts the contradictions.

[0080] 2. Ensure that the information extracted in the summary is clear and concise, only including the parts directly related to the key contradictions, and deleting irrelevant information.

[0081] 3. Process each data entry to generate a short summary, retaining the most critical information.

[0082] 4. Generate a final summary dataset, including all processed entries, arranged in order.

[0083] List of data entries:

[0084] - Data entry 1: <Data content 1>

[0085] - Data entry 2: <Data content 2> ......

[0086] Key contradictions:

[0087] -( ): <Core contradiction 1>

[0088] -( ): <Core contradiction 2>

[0089] Please output the summary of each data entry item by item, ensuring that each summary clearly indicates the relationship between the data and the contradictions, and output these summaries in order:}

[0090] The output of the large language model is the basis for the judgment summary;

[0091] Specifically, the rumor judgment and decision-making module is further configured to: in the fifth prompt word: require the large language model to analyze whether the information in the judgment basis summary is consistent with the information in the text to be detected, and judge whether to affirm or deny the core contradiction; require the large language model to judge the reliability of the corresponding judgment basis based on the information source of each judgment basis summary; require the large language model to judge whether the text to be detected is a rumor based on the reliability of the judgment basis summary and the impact on the core contradiction; require the large language model to output the analysis and judgment process of each judgment basis summary on the core contradiction of the text to be detected, and point out at least one of the most important judgment basis summaries in the analysis and judgment process.

[0092] In this embodiment, the fifth prompt word is:

[0093] {Now it is necessary to judge whether a piece of information is a rumor. The following are the judgment basis summaries from the rumor refutation knowledge base and the Internet , as well as the text to be detected that has been cleaned , as well as the core contradiction extracted from the text to be detected and . Please judge whether this information is a rumor based on these judgment basis summaries and the text to be detected. Please judge according to the following steps:

[0094] 1. Information integration: Comprehensively consider all relevant contents in the judgment basis summary , analyze whether they are consistent with the key information in the text to be detected , especially whether they support or refute the previously mentioned core contradiction and .

[0095] 2. Contradiction verification: Compare with the core contradiction and , judge whether these data generate new contradictions, and whether there are situations of information exaggeration, distortion or fabrication. If there are inconsistent or contradictory parts between the judgment basis summary and the text to be detected, focus on evaluating the impact of these parts on the rumor judgment.

[0096] 3. Information credibility: Based on the information source of each judgment basis summary, as well as the objectivity, authenticity, writing style, etc. of the text, analyze its impact on the rumor judgment. Give priority to information with high credibility.

[0097] 4. Final judgment: Based on the above analysis, judge whether this text is a rumor. Please give the final judgment according to the following criteria:

[0098] - If most of the summaries support the main content of the text to be detected and there are no obvious contradictions, it is determined as "not a rumor".

[0099] - If there is sufficient evidence indicating exaggeration, fabrication, distortion of facts, or information inconsistent with objective reality, it is determined as "rumor".

[0100] - If it is judged that the information in the abstract is inconsistent with the text to be detected, but the contradiction is not significant, it is determined as "possibly a rumor, further verification is required".

[0101] 5. Output of the analysis process:

[0102] - Please output the detailed analysis process, including the reference content of each judgment basis abstract, and indicate which judgment basis abstracts play a decisive role in the final judgment.

[0103] - Ensure to explain each judgment basis abstract referred to and how they support or refute the core contradiction in the text to be detected.

[0104] In this embodiment, it further includes an original text preprocessing module, which is used for: removing the HTML tags and comments in the original text ;

[0105] ;

[0106] Among them, represents the text after preprocessing and cleaning, defined as the text to be detected; represents removing HTML tags, and

[0107] is for removing comments. In this embodiment, these two methods are completed based on regular expressions to avoid the influence of HTML tags and comments in the subsequent detection process.

[0108] As a further improvement of this embodiment, the Internet search module is further used to: in the third prompt word, require the large language model to divide the credible search results into information without inspection value, information with inspection value and completely irrelevant information; the information without inspection value is an entity related to the core contradiction, but has no direct relationship with the core contradiction, or the information provided is useless, unreliable or irrelevant; the information with inspection value is related to the core contradiction and provides useful information or clues for further investigation; the completely irrelevant information is completely irrelevant to the entity mentioned in the core contradiction, or is completely irrelevant to the core contradiction, or is meaningless or misleading information;

[0109] In this embodiment, the Internet search module is also used to: establish a source database, and divide the sources into high-credibility sources and low-credibility sources; wherein, high-credibility sources have a certain degree of credibility, but may have some deviations and errors, and are usually well-known media or institutions in the industry, with a certain degree of credibility, including well-known mainstream media at home and abroad, reputable commercial websites and industry media, etc.; low-credibility sources lack credibility, and may be biased, misleading or exaggerated, including ordinary individuals on social media, forums of unknown origin, self-media accounts, etc. As a preferred embodiment of this embodiment, the high-credibility sources include media and commercial websites; the low-credibility sources include individuals, self-media and forums.

[0110] In the present embodiment, the large language models inputted by the first prompt word, the second prompt word, the third prompt word, the fourth prompt word, and the fifth prompt word may be the same or different large language models. As a preferred embodiment of the present embodiment, the first prompt word, the second prompt word, the third prompt word, the fourth prompt word, and the fifth prompt word are inputted into different large language models, which are selected based on their performance such as subdivision fields or reasoning or summarization, thereby providing stronger performance for rumor detection. In the present embodiment, the large language model may be one or more of ChatGPT, BERT, Claude, LLaMA, Grok, DeepSeek, Gemini, and Qwen, or other large language models.

[0111] This embodiment provides a rumor detection system based on a large language model, including a core contradiction extraction module, a fact-checking retrieval module, and a rumor judgment and decision-making module. The core contradiction extraction module is used to extract the core contradiction of entity relationships and the core contradiction of entity attributes from the text to be detected through the large language model. The fact-checking retrieval module includes a knowledge base retrieval module and an Internet retrieval module, which are used to retrieve relevant and reliable judgment bases from the knowledge base and the Internet based on the core contradiction of entity relationships and the core contradiction of entity attributes. The rumor judgment and decision-making module is used to analyze and judge the impact of the judgment bases on the authenticity of the text to be detected through the large language model based on the core contradiction and the judgment bases, and finally output the analysis and demonstration and the detection result based on the judgment bases. By extracting the core contradiction through the large language model, the present invention can reduce the probability of being interfered by irrelevant information, facilitate subsequent retrieval, improve the information utilization efficiency, and enhance the rumor detection accuracy. At the same time, by retrieving relevant judgment bases in the knowledge base and the Internet through the fact-checking retrieval module, the source of the judgment bases is enriched, and the hallucination of the large language model is avoided from affecting the detection result. Furthermore, the reliability of the judgment bases is further improved through the filtering of the information source database. Finally, the rumor judgment and decision-making module enables the large language model to analyze the impact of each judgment basis summary on the core contradiction, form a detection result, simulate the human judgment process, make the detection result based on the judgment basis, and be interpretable.

[0112] Embodiment 2:

[0113] Refer to Figure 2 As shown, this embodiment provides a rumor detection method based on a large language model that applies the system described in Embodiment 1, including:

[0114] Input the first prompt word and the text to be detected into the large language model to extract the core contradiction of entity relationships in the text to be detected. The core contradiction of entity relationships is the contradiction between the mutual relationships of multiple entities in the text to be detected output by the large language model and the facts. Input the second prompt word and the text to be detected into the large language model to extract the core contradiction of entity attributes in the text to be detected. The core contradiction of entity attributes is the contradiction between the attributes of the entities in the text to be detected output by the large language model and the facts.

[0115] Retrieve multiple pieces of knowledge base data most relevant to the core contradiction in the knowledge base through the RAG algorithm to obtain relevant knowledge base data. Retrieve the core contradiction from the Internet to obtain a retrieval result, filter the retrieval result through the information source database to obtain a reliable retrieval result. Input the third prompt word, the core contradiction, and the reliable retrieval result into the large language model to obtain a relevant reliable retrieval result. The core contradiction includes the core contradiction of entity relationships and the core contradiction of entity attributes.

[0116] Input the fourth prompt, the core contradiction, and the judgment basis into the large language model, summarize each judgment basis, and extract the content that mutually corroborates or contradicts the core contradiction to form a summary of the judgment basis; the judgment basis includes the relevant knowledge base data and the relevant credible search results; input the fifth prompt, the text to be detected, the core contradiction, and the summary of the judgment basis into the large language model, and analyze the impact of each summary of the judgment basis on the core contradiction to form a detection result.

[0117] Embodiment III:

[0118] This embodiment provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method described in Embodiment II.

[0119] Embodiment IV:

[0120] This embodiment provides a computer device, which includes a memory, a processor, and a program stored and executable on the memory. When the program is executed by the processor, it implements the steps described in Embodiment III.

[0121] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as these combinations of technical features do not conflict, they should be considered as within the scope described in this specification.

[0122] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.< / c>

Claims

1. A rumor detection system based on large language models, characterized in that, Including: A core contradiction extraction module, which is used to: input the first prompt word and the text to be detected into a large language model, and extract the core contradiction of entity relationships in the text to be detected. The core contradiction of entity relationships is the contradiction between the mutual relationships of multiple entities in the text to be detected output by the large language model and the facts; Input the second prompt word and the text to be detected into the large language model, and extract the core contradiction of entity attributes in the text to be detected. The core contradiction of entity attributes is the contradiction between the attributes of the entity in the text to be detected output by the large language model and the facts; A fact verification and retrieval module, including a knowledge base retrieval module and an Internet retrieval module; the knowledge base retrieval module is used to: retrieve multiple pieces of knowledge base data most relevant to the core contradiction in the knowledge base through the RAG algorithm to obtain relevant knowledge base data; the Internet retrieval module is used to: retrieve the core contradiction from the Internet to obtain a retrieval result, and filter the retrieval result through a source database to obtain a credible retrieval result; Input the third prompt word, the core contradiction, and the credible retrieval result into the large language model to obtain a relevant credible retrieval result; the core contradiction includes the core contradiction of entity relationships and the core contradiction of entity attributes; A rumor judgment and decision-making module, which is used to: input the fourth prompt word, the core contradiction, and the judgment basis into the large language model, summarize each judgment basis, and extract the content that mutually corroborates or contradicts the core contradiction therein to form a summary of the judgment basis; the judgment basis includes the relevant knowledge base data and the relevant credible retrieval result; input the fifth prompt word, the text to be detected, the core contradiction, and the summary of the judgment basis into the large language model, and analyze the influence of each summary of the judgment basis on the core contradiction to form a detection result.

2. The rumor detection system based on a large language model according to claim 1, wherein The core contradiction extraction module is further used to: add its time stamp when extracting the core contradiction of entity relationships in the text to be detected, and add its time stamp when extracting the core contradiction of entity attributes in the text to be detected; The fact verification and retrieval module is further used to: filter out the knowledge base data whose time is completely irrelevant to the corresponding time stamp when retrieving relevant knowledge base data; filter out the retrieval results whose time is completely irrelevant to the corresponding time stamp when retrieving the core contradiction from the Internet.

3. A rumor detection system based on a large language model according to claim 1, characterized in that, The knowledge base retrieval module is further used to: crawl the data of the rumor refutation platform and establish a knowledge base.

4. A rumor detection system based on a large language model according to claim 1, characterized in that, The knowledge base retrieval module is further used to: Vectorize each piece of knowledge base data in the knowledge base to form vectorized knowledge base data; vectorize the core contradiction of entity relationships and the core contradiction of entity attributes to form vectorized core contradiction of entity relationships and vectorized core contradiction of entity attributes; Calculate the similarity score between each vectorized core contradiction of entity relationships and each vectorized knowledge base data: ; Among them, represents the core contradiction of the vectorized entity relationship, represents the th vectorized knowledge base data; Calculate the similarity score between each vectorized core contradiction of entity attributes and each vectorized knowledge base data: ; Among them, represents the core contradiction of the vectorized entity attributes, represents the th vectorized knowledge base data; Take the larger value between the similarity score of the vectorized entity relationship core contradiction and the vectorized knowledge base data and the similarity score of the vectorized entity attribute core contradiction and the vectorized knowledge base data as the final similarity score of the corresponding knowledge base data: ; Select multiple pieces of knowledge base data with the highest final similarity scores as the relevant knowledge base data.

5. A rumor detection system based on a large language model according to claim 1, characterized in that, The Internet retrieval module is further configured to: require the large language model in the third prompt word to classify the credible retrieval results into information with no inspection value, information with inspection value, and completely irrelevant information; the information with no inspection value involves entities related to the core contradiction but has no direct relationship with the core contradiction, or the information provided is useless, untrustworthy, or irrelevant; the information with inspection value is related to the core contradiction and provides useful information or clues for further investigation; the completely irrelevant information completely does not involve the entities mentioned in the core contradiction, or is completely irrelevant to the core contradiction, or is meaningless or misleading information.

6. The rumor detection system based on a large language model according to claim 1, wherein, The Internet retrieval module is further configured to: establish a source database, and classify the sources into high-credibility sources and low-credibility sources; the high-credibility sources include media and commercial websites; the low-credibility sources include individuals, self-media, and forums.

7. A rumor detection system based on a large language model according to claim 1, characterized in that, The rumor judgment and decision-making module is further configured to: in the fifth prompt word: require the large language model to analyze whether the information in the judgment basis summary is consistent with the information in the text to be detected, and judge whether to affirm or deny the core contradiction; Require the large language model to judge the reliability of the corresponding judgment basis based on the source of each judgment basis summary; Require the large language model to judge whether the text to be detected is a rumor based on the reliability of the judgment basis summary and the impact on the core contradiction; require the large language model to output the analysis and judgment process of each judgment basis summary on the core contradiction of the text to be detected, and point out at least one of the most important judgment basis summaries in the analysis and judgment process.

8. A large language model-based rumor detection method applied to the system according to any one of claims 1 to 7, characterized in that, Include: Input the first prompt word and the text to be detected into the large language model, and extract the core contradiction of the entity relationship in the text to be detected. The core contradiction of the entity relationship is the contradiction between the mutual relationship of multiple entities in the text to be detected output by the large language model and the fact; Input the second prompt word and the text to be detected into the large language model, and extract the core contradiction of the entity attribute in the text to be detected. The core contradiction of the entity attribute is the contradiction between the attribute of the entity in the text to be detected output by the large language model and the fact; Retrieve multiple pieces of knowledge base data most relevant to the core contradiction in the knowledge base through the RAG algorithm to obtain relevant knowledge base data; Retrieve the core contradiction from the Internet to obtain retrieval results, and filter the retrieval results through the source database to obtain credible retrieval results; Input the third prompt word, the core contradiction, and the credible retrieval results into the large language model to obtain relevant credible retrieval results; The core contradiction includes the core contradiction of the entity relationship and the core contradiction of the entity attribute; Input the fourth prompt word, the core contradiction, and the judgment basis into the large language model. Summarize each judgment basis and extract the content that mutually corroborates or contradicts the core contradiction to form a judgment basis summary. The judgment basis includes the relevant knowledge base data and the relevant credible retrieval results. Input the fifth prompt word, the text to be detected, the core contradiction, and the judgment basis summary into the large language model, and analyze the impact of each judgment basis summary on the core contradiction to form a detection result.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method according to claim 8.

10. A computer device, characterized in that, The computer device includes a memory, a processor, and a program stored and executable on the memory. When the program is executed by the processor, it implements the steps of the method according to claim 8.

Citation Information

Patent Citations

  • Government affair processing method of big language model based on knowledge graph and electronic equipment

    CN119782545A

  • Information sending method and apparatus based on rumor prediction model, and computer device

    WO2022001517A1

Cited By

  • Misleading information detection method and system based on artificial intelligence

    CN122364376A