Rumor recognition method and system based on fact description and expansion graph
Through the method based on fact description and augmented graphs, a fact network of event-type information is constructed and graph classification model training is carried out, which solves the problem of rumors recognition in newly-released event-type information, and achieves accurate rumors recognition and domain expansion.
Patent Information
- Application Number
- CN202510056727.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-09
AI Technical Summary
It is difficult for the prior art to effectively analyze and process rumors in newly-occurred incident information online, especially when the event occurs locality and the information-related person and matters are particular.
Using a method based on fact description and augmented graph, we use the graph classification model to classify rumors and non-rumor information by obtaining event-type information, extracting facts, and constructing fact description and augmented graphs.
It realizes accurate rumor identification of newly-occurred incident-type information, has good domain expansion, can be used to judge the authenticity of information explosive dissemination in different fields in the early stages of explosive dissemination, and supplements the rumor identification method based on the communication structure and text style.
Smart Images

Figure CN119961410A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer content security, and in particular to a rumor identification method and system based on fact description and expansion graph. Background Art
[0003] Among the existing rumor identification technologies and methods, most of them are based on intelligent classification based on the content of the rumor itself and the content of interactive information during the spread of rumors on social media, for example, analyzing language style, stance or emotion. And based on the structure of information dissemination, the rumor is distinguished according to its special characteristics in dissemination. The most reliable way to confirm rumors is fact-checking, but fact-checking is a difficult point in the intelligent rumor identification method, and there are few methods directly targeting fact-checking. Many rumors on social media are not transactional information, but event information. Transactional information can be verified through retrieval and other methods, but event information needs to be verified on the spot.
[0004] Patent documents CN202410144684.6 and CN202410514089.7 use data enhancement and few-sample algorithms to achieve fact verification, and CN202410796635.0 uses comments and forwarding information to strengthen the integration of external knowledge. This verification method is more suitable for the verification of transactional and common sense information. Patent document CN202311689864.4 Chinese invention patent discloses a rumor detection method based on a multi-level untrue propagation structure. First, a large-scale propagation structure data collection is carried out in social media, and the information text and its propagation structure are modeled, and a classification algorithm is used to identify rumors. In order to solve the problem of the difference between the sample set and the data to be detected in the rumor identification method when the information topic changes, the Chinese invention patent with the patent document publication number CN202410357819.7 aims to improve the cross-domain ability of the recognition model for different topics, and uses domain adversarial learning to realize the adaptive recognition method of rumors in new fields.
[0005] Rumors spread faster than normal information on social media, and the identification of rumors requires accurate and rapid responses. In addition, event-type information, where events occur geographically locally and the people and events involved in the information are specific, the economic and timeliness of fact-checking new events is poor.
[0006] However, the above-mentioned patent documents are unable to perform online analysis and processing of new event-type information. Summary of the invention
[0007] In view of the defects in the prior art, the purpose of the present invention is to provide a rumor identification method and system based on fact description and expansion graph.
[0008] A rumor identification method based on fact description and expansion graph provided by the present invention includes:
[0009] Step S1: Acquire event information and extract facts from the event information; the event information includes original information, subsequent information and interactive information;
[0010] Step S2: aligning the facts and constructing an expansion graph to obtain a fact description and an expansion graph for rumor analysis;
[0011] Step S3: Using the fact description and the expanded graph as samples, a graph classification model is trained to ultimately achieve the classification of rumor and non-rumor information.
[0012] Preferably, the facts in the event-type information description include the entities mentioned in the information and the behaviors, relationships between the entities, and spatiotemporal data;
[0013] The entities include people, things, and objects; the spatiotemporal data include the time and place of the event;
[0014] The original information is the information that triggers a large spread of an event, the subsequent information is the subsequent supplement to the event-type information by the publisher of the original information, and the interactive information is the comments and replies in the original information and subsequent information.
[0015] Preferably, the step S1 comprises:
[0016] Step S1.1: Classify the sentences of event-type information, remove interrogative sentences and speculative sentences, and filter out descriptive sentences;
[0017] Step S1.2: performing named entity recognition on the descriptive sentence to obtain existing data and missing data corresponding to the entity;
[0018] Step S1.3: extracting event elements in the descriptive sentence based on the existing data and missing data, that is, the relationship between entities; the relationship between entities includes the relationship between time and action behavior, more attributes of people or things, and the relationship between people or things.
[0019] Preferably, the step S2 comprises:
[0020] Step S2.1: Taking the extracted original information and the facts contained in the original information as roots, a fact network is constructed, wherein the original information is the initial node of the fact network and the corresponding publishing timestamp is an attribute;
[0021] Step S2.2: Expand the facts extracted from the subsequent information and interactive information to the existing fact network, construct a fact expansion network, and obtain a fact description and an expansion graph; each information has a corresponding information node, the subsequent information node is directly connected to the previous information of the current subsequent information node, and the interactive information is connected to the information corresponding to the current interactive information; when expanding facts, they are aligned through the character behavior nodes, and the new nodes are connected to the aligned existing nodes according to the data relationship;
[0022] Step S2.3: Analyze the frequency statistics of questionable mentions of facts in all information and add them to the fact description and expansion diagram;
[0023] Step S2.4: Based on the fact description and the expanded graph described in step S2.3, simplify the network;
[0024] The simplification includes deleting information nodes that are not connected to fact nodes, directly connecting corresponding subsequent information nodes to previous information nodes; and merging repeated fact nodes in subsequent information and interactive information.
[0025] Preferably, step S3 comprises:
[0026] Step S3.1: Collect information and construct a current fact description and an expanded graph, mark rumor information and non-rumor information, and use the graph features to apply a graph classification algorithm to train the model;
[0027] The features of the graph include data such as fact structure and fact mentions;
[0028] Step S3.2: Use the trained classification model to classify the information.
[0029] According to the present invention, a rumor identification system based on fact description and expansion graph is provided, comprising:
[0030] Module M1: Acquire event information and extract facts from the event information; the event information includes original information, subsequent information and interactive information;
[0031] Module M2: aligning the facts and constructing an expansion graph to obtain a fact description and an expansion graph for rumor analysis;
[0032] Module M3: Using the fact description and the expanded graph as samples, the graph classification model is trained to ultimately achieve the classification of rumor and non-rumor information.
[0033] Preferably, the facts in the event-type information description include the entities mentioned in the information and the behaviors, relationships between the entities, and spatiotemporal data;
[0034] The entities include people, things, and objects; the spatiotemporal data include the time and place of the event;
[0035] The original information is the information that triggers a large spread of an event, the subsequent information is the subsequent supplement to the event-type information by the publisher of the original information, and the interactive information is the comments and replies in the original information and subsequent information.
[0036] Preferably, the module M1 comprises:
[0037] Module M1.1: Classify the sentences of event-type information, remove interrogative sentences and speculative sentences, and filter out descriptive sentences;
[0038] Module M1.2: Perform named entity recognition on the descriptive sentence to obtain existing data and missing data corresponding to the entity;
[0039] Module M1.3: Extract the event elements in the descriptive sentence based on the existing data and missing data, that is, the relationship between entities; the relationship between entities includes the relationship between time and action behavior, more attributes of people or things, and the relationship between people or things.
[0040] Preferably, the module M2 comprises:
[0041] Module M2.1: Taking the extracted original information and the facts contained in the original information as the root, construct a fact network, wherein the original information is the initial node of the fact network and the corresponding release timestamp is the attribute;
[0042] Module M2.2: Expand the facts extracted from the subsequent information and interactive information to the existing fact network, build a fact expansion network, and obtain the fact description and expansion graph; each information has a corresponding information node, the subsequent information node is directly connected to the previous information of the current subsequent information node, and the interactive information is connected to the information corresponding to the current interactive information; when expanding facts, they are aligned through the character behavior nodes, and the new nodes are connected to the aligned existing nodes according to the data relationship;
[0043] Module M2.3: Analyze the frequency statistics of questionable mentions of facts in all information and add them to the description of the facts and the expansion diagram;
[0044] Module M2.4: Simplify the network based on the fact description and expanded diagram described in Module M2.3;
[0045] The simplification includes deleting information nodes that are not connected to fact nodes, directly connecting corresponding subsequent information nodes to previous information nodes; and merging repeated fact nodes in subsequent information and interactive information.
[0046] Preferably, the module M3 comprises:
[0047] Module M3.1: Collect information and construct a description of current facts and an expanded graph, label rumor information and non-rumor information, and use graph features to apply graph classification algorithms to train models;
[0048] The features of the graph include data such as fact structure and fact mentions;
[0049] Module M3.2: Use the trained classification model to classify information.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] 1. The present invention targets newly-emerged event-type information. With the influx of subsequent information and interactive information in the information dissemination process, the present invention adopts fact extraction, alignment, and expansion methods to construct fact alignment and expansion graphs of the information, which is suitable for classifying rumors for newly-emerged event-type information.
[0052] 2. The present invention has good scalability in fields and has little to do with the subject field of information. It can be applied to various fields such as social news, industrial economic news, etc. It can judge the authenticity of information through the fact expansion model in the early stage of explosive dissemination of information.
[0053] 3. Compared with existing rumor identification based on communication structure, text style, etc., the present invention identifies from the perspective of factual relationships and can serve as a supplement to existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0055] Figure 1 It is a schematic diagram of the working method of the present invention;
[0056] Figure 2 This is an example diagram of the construction of the fact alignment and expansion graph, where the left picture is before simplification and the right picture is after simplification;
[0057] Figure 3 This is a diagram of the rumor detection classification architecture based on a two-layer stacked graph attention model in an embodiment. DETAILED DESCRIPTION
[0058] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0059] Example 1
[0060] According to the rumor identification method based on fact description and expansion graph provided by the present invention, Figures 1 to 3 As shown, including:
[0061] Step S1: Obtain event information and extract facts from the information. Facts are extracted from event information by means of named entity recognition, entity relationship recognition, etc. The event information includes original information, follow-up information, and interactive information. The facts described in the event information include the entities mentioned in the information and the behaviors and relationships between the entities, as well as spatiotemporal data. The entities include people, things, and objects, and the spatiotemporal data includes the time and place of the event.
[0062] The output data involved in step S1 is stored in the fact storage and retrieval library. Step S1 includes:
[0063] Step S1.1: Classify the sentences in the information, remove interrogative sentences and speculative sentences, and filter out descriptive sentences.
[0064] Step S1.2: Perform named entity recognition on the descriptive sentence. Entity types include event time, location, related people or animals or things, actions and behaviors. The existing data and missing data in the four elements of time, location, people and behavior of the event are saved and marked. Missing data is stored as temporarily unknown according to the content.
[0065] Step S1.3: extracting the event elements in the descriptive sentence, that is, the relationship between entities. The relationship between entities includes the relationship between time and action, more attributes of people or things, and the relationship between people or things. For example, the social relationship of people, the subordinate relationship between organization names and people, the relationship between actions and locations, and the time relationship between actions.
[0066] Step S2: Align the facts and construct an expansion graph to obtain a fact description and expansion graph for rumor analysis. A fact network is established based on the facts in the original information, and the facts extracted from the subsequent information and interactive information are aligned with the facts of the original information, and the fact network is expanded and simplified, finally forming a fact description and expansion graph that can be used for rumor analysis.
[0067] The fact description and supplementary graph refers to a graph composed of nodes or connections based on the facts described in the original information and subsequent information released by the original information publisher, as well as all the interactive information of these information. Subsequent information is the subsequent supplement of the original information publisher to the event-type information, and interactive information is the comments and replies in the original information and subsequent information. Starting with the establishment of the first node with the original information, the subsequent information and interactive information establish nodes and connect them according to the time sequence and interaction order. The facts extracted from each piece of information establish nodes respectively and connect to the information nodes. The facts in the subsequent and interactive information need to be aligned with the facts in the original information, and the aligned original information fact nodes are connected to the fact nodes of the subsequent and interactive information. Specifically, the step S2 includes:
[0068] Step S2.1: Take the extracted original information and the facts contained in the original information as the root to construct a fact network, where the information is the initial node of the network, the information release timestamp is the attribute, the time of the event, the place of the event, the person, and the behavior description extracted from the information are the direct connection nodes of the information node, and the relationship between the person, behavior and things is the node connection. The fact node of the original information allows all attributes to be unknown. The original information is the information that triggers a large spread of an event, not the earliest published information about the event. The earliest published information is added to the information flow during the spread process, which is the subsequent information or interactive information.
[0069] Step S2.2: Expand the facts extracted from subsequent information and interactive information such as reply comments to the existing fact network to build a fact expansion network. Subsequent information nodes are directly connected to their predecessor information, and interactive information is connected to the information of its reply or comment. When expanding facts, they are aligned through character behavior nodes, that is, the differences in character titles are processed, and different character titles and specific characters with missing titles are merged. After the nodes are aligned, the new nodes are connected to the aligned existing nodes according to the data relationship. Nodes with unknown attributes are not allowed to exist in the nodes of the expanded facts.
[0070] Step S2.3: Analyze the frequency statistics of questionable mentions of facts in all interactive information and add them to the fact description and expansion diagram.
[0071] Step S2.4: Based on the fact description and expansion graph described in step S2.3, simplify the network. If there is no fact node or connection connected to the subsequent information and interactive information node, that is, the node has no supplementary facts, merge the node into the previous node. If there are repeated facts in the subsequent information and interactive information nodes, merge the repeated fact nodes.
[0072] Step S3: Using the fact description and the expanded graph as samples, a graph classification model is trained to finally classify rumor and non-rumor information. Step S3 includes:
[0073] Step S3.1: Classification model training. Collect information and build a current fact description and expansion graph, manually mark rumor information and non-rumor information, and use the graph features to apply the graph classification algorithm to train the model. The graph features include fact structure, fact mentions and other data.
[0074] Step S3.2: Use the trained classification model to classify the information.
[0075] Further, in conjunction with the accompanying drawings, the rumor identification method based on fact description and expansion graph of the present invention is specifically described as follows:
[0076] like Figure 1 As shown, this embodiment relates to a method and system for extracting facts from new event information in social media, constructing fact descriptions and expansion graphs, and then using graph classification technology to realize rumor identification. The system proposed by the present invention does not depend on the type of social media where the information is published. The data detected in this embodiment is collected from the Twitter website. An original tweet is selected as the original information, and the forwarding and reply information of the original tweet that is not the author of the original tweet is used as the interactive information. The subsequent remarks of the original author in the forwarding and reply information are used as the subsequent information, and the original texts of these information are comprehensively collected. The detailed process of each step in the embodiment will be explained below.
[0077] 1. Information Extraction of Event-based Information
[0078] Step 1.1: Classify the sentences in the information, remove interrogative sentences and speculative sentences, and filter out descriptive sentences. Here, a large language model is used, and the prompt word engineering method is used to instruct the large model to remove speculative and interrogative sentences in the information text, and only save and output the affirmative descriptive sentences in the original information.
[0079] Step 1.2: Briefly describe the event for the descriptive sentences in the information. Here, a large language model prompting project is used. First, the information is summarized to form a brief description of the event. The brief description of the event is the behavior of a person / institution / event, such as a company reducing the price of a product. If there are multiple events in the information, the event with the most information content is selected. Secondly, named entity recognition is performed on the sentences related to the brief description of the event, and three types of entities {'time', 'place', 'person / institution / thing'} are extracted. The number of these three types of entities can be more than one. If there is no specific entity in the information, such as no time, the attribute of time is stored as unknown. For the 'person / institution / thing' entity, more attributes are extracted, such as identifying the occupation, age, clothing appearance, name for the person, identifying the field, scale, region for the organization, and identifying the size, color, and status for the thing.
[0080] Step 1.3: Extract the event elements in the descriptive sentences in the information, that is, the relationship between entities. Use the large language model to confirm the relationship between the extracted entities and the event summary. Entities that are directly related to the event summary are saved, otherwise they are deleted. When confirming subsequent information and interactive information in this step, the previous information of these information, the author information of the information, etc., need to be submitted to the large language model for judgment as a basis.
[0081] 2. Fact Description and Expansion Graph Construction
[0082] Step 2.1: Take the original information extracted in steps 1.2 and 1.3 and the facts they contain as roots to construct a fact network, which is a directed acyclic graph. Figure 2 As shown in the left figure, information 1 is the original information, which is the initial node of the network, the subsequent information directly connected to it, and the event entity extracted from the information, which contains entities with empty attributes.
[0083] Step 2.2: Expand the facts extracted from the subsequent information and interactive information such as reply comments in steps 1.2 and 1.3 to the existing fact network to build a fact description and expansion network. The subsequent information node is directly connected to its previous information, and the interactive information is connected to the information it replies to or comments on. The entity nodes extracted from the information are connected to the information they belong to. Entities with empty attributes in this step are not included in the graph. The newly added entity nodes will align the entity nodes connected to the previous information nodes. The alignment algorithm uses the large language model confirmation method, with the same type of nodes in the existing network and their belonging information as the basis. If two entity nodes belong to the same entity, an alignment connection is established.
[0084] Step 2.3: Analyze the interrogative sentences and questioning sentences in all subsequent information interactions for mentions of fact nodes contained in all previous information. The number of mentions is saved as the attribute value of the nearest fact node, for example Figure 2 If there is a mention of a place in information 4 in the left figure, the frequency of the mention will be saved in the place node of information 3. If there is a mention of time, the frequency of the mention will be saved in the time node connected to information 1.
[0085] Step 2.4: Simplify the network. Figure 2 As shown in the right figure, simplification has two steps: 1. Merging aligned entities. The factual entity node Person 3 of information 4 and the node Person 3 of information 3 are in an aligned relationship, and the person attribute is an inclusion relationship. Therefore, the Person 3 node of information 4 as a subsequent node is merged into the Person 3 node of information 3, and the mention and other data are added, and all attributes are retained; 2. Simplification of information nodes, deleting information nodes that are not connected to entity nodes. Figure 2The information 2 node in the left figure has no fact node connected to it, but is connected to information 1 and information 3. It is deleted in the simplification, and information 3 will be directly connected to information 1, such as Figure 2 Right picture.
[0086] 3. Rumor Identification Based on Fact Description and Extended Graph
[0087] The true and false labels of the original information in the sample set are manually annotated, and the graph attention network model is used to perform classification prediction of the original information nodes in the graph. The training of the graph attention network is completed by reducing the sum of the losses of the training set.
[0088] The fact description and expansion graph includes two types of nodes: information nodes and fact entity nodes. Fact entity nodes are divided into three categories: time, place, and person. The feature vector of the fact entity node consists of [type, mention, number of attributes], where the type is encoded using one-hot. The feature vector of the information node is initialized randomly. The feature vectors of the information node and the fact entity node form a node feature matrix as the input of the graph algorithm. Generate the node adjacency matrix A based on the graph.
[0089] The node feature transformation formula of the graph attention network is: h' v =Wh v , where W is the weight matrix of the graph attention layer, h v is the node feature matrix. v Substitute the following formula:
[0090] e ij =LeakyReLU(a T (h' i ||h' j ))
[0091] Among them, α is the attention vector to be trained, and j represents the adjacent connection node of node i.
[0092] E i Substitute j into the following formula:
[0093]
[0094] Among them, N i is the set of neighbor nodes of node i, α ij is the calculated attention weight of the neighboring nodes.
[0095] α ij Substitute the formula to get the attention output
[0096] like Figure 3 As shown in the figure, the algorithm uses two layers of attention stacking, and the node feature vector z of the original information node in the last layer is1 , where, in order to distinguish different attention layers, the input layer node feature vector is X i , the output layer node feature vector is Z i . Use the fully connected layer to calculate the output y = w out z 1 +b, where W and b are all fully connected parameters. Substitute y into the Sigmoid activation function to obtain the prediction result. Use binary cross entropy to calculate the loss Loss = ylog(y′) + (1-y)log(1-y′) and perform back propagation training on the network.
[0097] The above-mentioned embodiments and results show that the method and system of the present invention establish a fact description graph of a certain original information, and use the subsequent information of the information and interactive information to expand the facts. In the field of rumor detection, the facts, mentions and their associations in the application information are used to establish a fact description and expansion graph of the information, thereby realizing rumor detection of event-type information based on factual content. In the embodiment, an accuracy rate of 85% was achieved. In summary, the system described in the present invention can eliminate the interference of existing subjective factors such as communication structure, stylistic style, and language style in a short period of time when the fact verification of event-type information is difficult, and provides a new detection idea.
[0098] Example 2
[0099] The present invention also provides a rumor identification system based on factual description and expanded graph. The rumor identification system based on factual description and expanded graph can be implemented by executing the process steps of the rumor identification method based on factual description and expanded graph, that is, those skilled in the art can understand the rumor identification method based on factual description and expanded graph as a preferred implementation of the rumor identification system based on factual description and expanded graph.
[0100] According to the present invention, a rumor identification system based on fact description and expansion graph is provided, comprising:
[0101] Module M1: Acquire event-type information and extract facts from the event-type information. The event-type information includes original information, subsequent information and interactive information. The facts described in the event-type information include the entities mentioned in the information and the behaviors, relationships between the entities, and spatiotemporal data. The entities include people, events and things. The spatiotemporal data include the time and place of the event. The original information is the information that triggers the widespread dissemination of an event, the subsequent information is the subsequent supplement of the event-type information by the publisher of the original information, and the interactive information is the comments and replies in the original information and subsequent information. The module M1 includes: Module M1.1: Classify the sentences of the event-type information, remove interrogative sentences and speculative sentences, and filter out descriptive sentences. Module M1.2: Perform named entity recognition on the descriptive sentences to obtain the existing data and missing data corresponding to the entities. Module M1.3: Extract the event elements in the descriptive sentences, that is, the relationships between entities, based on the existing data and missing data. The relationships between entities include the relationship between time and action behavior, more attributes of people or things, and the relationship between people or things.
[0102] Module M2: Align the facts and construct an expansion graph to obtain a fact description and an expansion graph for rumor analysis. The module M2 includes: Module M2.1: Take the extracted original information and the facts contained in the original information as the root to construct a fact network, wherein the original information is the initial node of the fact network and the corresponding release timestamp is the attribute. Module M2.2: Expand the facts extracted from the subsequent information and interactive information to the existing fact network, construct a fact expansion network, and obtain a fact description and an expansion graph. Each information has a corresponding information node, the subsequent information node is directly connected to the previous information of the current subsequent information node, and the interactive information is connected to the information corresponding to the current interactive information. When expanding facts, alignment is performed through the character behavior nodes, and the new nodes are connected to the aligned existing nodes according to the data relationship. Module M2.3: Analyze the frequency statistics of questionable mentions of facts in all information and add them to the fact description and expansion graph. Module M2.4: Based on the fact description and expansion graph described in module M2.3, network simplification is performed. The simplification includes deleting information nodes that are not connected to fact nodes, directly connecting corresponding subsequent information nodes to previous information nodes, and merging repeated fact nodes in subsequent information and interactive information.
[0103] Module M3: Using the fact description and extended graph as samples, the graph classification model is trained to finally classify rumor and non-rumor information. The module M3 includes: Module M3.1: Collect information and form the current fact description and extended graph, mark rumor information and non-rumor information, and use the graph features to apply the graph classification algorithm to train the model. The graph features include fact structure, fact mention and other data. Module M3.2: Use the trained classification model to classify the information.
[0104] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.
[0105] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A rumor identification method based on fact description and extended graph, characterized in that: include: Step S1: Acquire event information and extract facts from the event information; The event-type information includes original information, subsequent information and interactive information; Step S2: aligning the facts and constructing an expansion graph to obtain a fact description and an expansion graph for rumor analysis; Step S3: Using the fact description and the expanded graph as samples, a graph classification model is trained to ultimately achieve the classification of rumor and non-rumor information.
2. The rumor identification method based on fact description and extended graph according to claim 1 is characterized in that: The facts in event-type information description include the entities mentioned in the information and the behaviors and relationships between the entities, as well as spatiotemporal data; The entities include people, things, and objects; the spatiotemporal data include the time and place of the event; The original information is the information that triggers a large spread of an event, the subsequent information is the subsequent supplement to the event-type information by the publisher of the original information, and the interactive information is the comments and replies in the original information and subsequent information.
3. The rumor identification method based on fact description and extended graph according to claim 2 is characterized in that: The step S1 comprises: Step S1.1: Classify the sentences of event-type information, remove interrogative sentences and speculative sentences, and filter out descriptive sentences; Step S1.2: performing named entity recognition on the descriptive sentence to obtain existing data and missing data corresponding to the entity; Step S1.3: extracting event elements in the descriptive sentence based on the existing data and missing data, that is, the relationship between entities; the relationship between entities includes the relationship between time and action behavior, more attributes of people or things, and the relationship between people or things.
4. The rumor identification method based on fact description and extended graph according to claim 1 is characterized in that: The step S2 comprises: Step S2.1: Taking the extracted original information and the facts contained in the original information as roots, a fact network is constructed, wherein the original information is the initial node of the fact network and the corresponding publishing timestamp is an attribute; Step S2.2: Expand the facts extracted from the subsequent information and interactive information to the existing fact network, construct a fact expansion network, and obtain a fact description and an expansion graph; each information has a corresponding information node, the subsequent information node is directly connected to the previous information of the current subsequent information node, and the interactive information is connected to the information corresponding to the current interactive information; when expanding facts, they are aligned through the character behavior nodes, and the new nodes are connected to the aligned existing nodes according to the data relationship; Step S2.3: Analyze the frequency statistics of questionable mentions of facts in all information and add them to the fact description and expansion diagram; Step S2.4: Based on the fact description and the expanded graph described in step S2.3, simplify the network; The simplification includes deleting information nodes that are not connected to fact nodes, directly connecting corresponding subsequent information nodes to previous information nodes; and merging repeated fact nodes in subsequent information and interactive information.
5. The rumor identification method based on fact description and extended graph according to claim 1 is characterized in that: The step S3 comprises: Step S3.1: Collect information and construct a current fact description and an expanded graph, mark rumor information and non-rumor information, and use the graph features to apply a graph classification algorithm to train the model; The features of the graph include data such as fact structure and fact mentions; Step S3.2: Use the trained classification model to classify the information.
6. A rumor identification system based on fact description and extended graph, characterized in that: include: Module M1: obtaining event information and extracting facts from the event information; The event-type information includes original information, subsequent information and interactive information; Module M2: aligning the facts and constructing an expansion graph to obtain a fact description and an expansion graph for rumor analysis; Module M3: Using the fact description and the expanded graph as samples, the graph classification model is trained to ultimately achieve the classification of rumor and non-rumor information.
7. The rumor identification system based on fact description and extended graph according to claim 6 is characterized in that: The facts in event-type information description include the entities mentioned in the information and the behaviors and relationships between the entities, as well as spatiotemporal data; The entities include people, things, and objects; the spatiotemporal data include the time and place of the event; The original information is the information that triggers a large spread of an event, the subsequent information is the subsequent supplement to the event-type information by the publisher of the original information, and the interactive information is the comments and replies in the original information and subsequent information.
8. The rumor identification system based on fact description and extended graph according to claim 7 is characterized in that: The module M1 comprises: Module M1.1: Classify the sentences of event-type information, remove interrogative sentences and speculative sentences, and filter out descriptive sentences; Module M1.2: Perform named entity recognition on the descriptive sentence to obtain existing data and missing data corresponding to the entity; Module M1.3: Extract the event elements in the descriptive sentence based on the existing data and missing data, that is, the relationship between entities; the relationship between entities includes the relationship between time and action behavior, more attributes of people or things, and the relationship between people or things.
9. The rumor identification system based on fact description and extended graph according to claim 6, characterized in that: The module M2 comprises: Module M2.1: Taking the extracted original information and the facts contained in the original information as the root, construct a fact network, wherein the original information is the initial node of the fact network and the corresponding release timestamp is the attribute; Module M2.2: Expand the facts extracted from the subsequent information and interactive information to the existing fact network, build a fact expansion network, and obtain the fact description and expansion graph; each information has a corresponding information node, the subsequent information node is directly connected to the previous information of the current subsequent information node, and the interactive information is connected to the information corresponding to the current interactive information; when expanding facts, they are aligned through the character behavior nodes, and the new nodes are connected to the aligned existing nodes according to the data relationship; Module M2.3: Analyze the frequency statistics of questionable mentions of facts in all information and add them to the description of the facts and the expansion diagram; Module M2.4: Simplify the network based on the fact description and expanded diagram described in Module M2.3; The simplification includes deleting information nodes that are not connected to fact nodes, directly connecting corresponding subsequent information nodes to previous information nodes; and merging repeated fact nodes in subsequent information and interactive information.
10. The rumor identification system based on fact description and extended graph according to claim 6, characterized in that: The module M3 comprises: Module M3.1: Collect information and construct a description of current facts and an expanded graph, label rumor information and non-rumor information, and use graph features to apply graph classification algorithms to train models; The features of the graph include data such as fact structure and fact mentions; Module M3.2: Use the trained classification model to classify information.
Citation Information
Patent Citations
Rumor detection method based on multi-level inauthenticity propagation structure
CN117852532A
Construction method of cross-domain rumor detection model, storage medium and terminal equipment
CN118095291A
Few-sample supervision Chinese fact checking system enhanced by rumor detection data
CN118133151A
Rumor detection method based on bidirectional graph neural network
CN118349681A
Rumor detection method based on node chained semantic features and knowledge fusion
CN118377918A