A Motif-based Heterogeneous Graph Neural Network Algorithm for Fake News Detection

By constructing news heterogeneous pictures and capturing higher-order semantic patterns in social networks using attention mechanisms, the problem of suboptimal performance of existing fake news detection algorithms is solved, and more efficient fake news detection is achieved.

CN115269853BActive Publication Date: 2025-07-18DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210949264.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-07-18
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

Existing fake news detection algorithms ignore higher-order semantic patterns in social networks, resulting in suboptimal detection performance.

Method used

By constructing news heterogeneous graphs, extracting third-order heterogeneous model instances, and capturing key information using instance-level and semantic-level attention mechanisms to learn effective news representations.

Benefits of technology

Improve the accuracy and performance of fake news detection, and solve the problem of suboptimal detection caused by ignoring higher-order semantic structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115269853B_ABST
    Figure CN115269853B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of fake news detection and provides a fake news detection algorithm based on motif-based heterogeneous graph neural network. First, the original data of social media is constructed into a news heterogeneous graph; secondly, all types of nodes are mapped to the same feature space, and instances are extracted according to each heterogeneous motif type respectively; then, the instance-level attention mechanism is used to aggregate all motif instances of the same type into the corresponding news nodes to capture key instance information; then, for different types of heterogeneous motifs, the semantic-level attention mechanism is used to adaptively aggregate different news semantic embeddings; finally, the representation of the news is used for the downstream fake news detection task. The present invention takes into account the large number of heterogeneous high-order patterns existing in social platforms, and through a two-layer attention mechanism, learns efficient news node representations and improves the effect of fake news detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network representation learning, and relates to a motif-based heterogeneous graph neural network fake news detection algorithm, which can be used for fake news detection in social media. Background Art

[0002] News has rich content and has long been an important source of information for people, influencing people's decisions in different activities. However, most fake news has attractive headlines and misleading content, attracting more readers and leading them to make wrong choices. In addition, the Internet has greatly accelerated the spread of information, and the threat of fake news has become more serious. Currently, fake news detection is a promising research direction, and in the modern social network environment, there is an urgent need for effective detection algorithms.

[0003] The current mainstream fake news detection algorithms on social media can be classified into the following categories:

[0004] Fake news detection algorithms based on news content. The work "FakeNews Detection Using Multiple-View Text Representation" published by Ha and Gao in PRICAI in 2021 proposed WES, which extracts three different representation views (word-level, sentence-level, and sentiment features) in the text content to classify news articles. The work "Multi-Modal fake news Detection on Social Media with Dual Attention FusionNetworks" published by Yang et al. in the IEEE Symposium on Computers and Communications in 2021 proposed DAFN, which fuses information from text and image modalities and uses BERT embeddings to obtain multi-modal fusion representations for fake news detection. The work "SAFE: Similarity-AwareMulti-modal Fake News Detection" published by Zhou et al. in PAKDD in 2020 proposed the SAFE method, which calculates the computational similarity between text and visual information and then jointly learns their representations to detect fake news. The work "FakeNews Detection via Knowledge-driven Multimodal Graph Convolutional Networks" published by Wang et al. in ICMR in 2020 proposed KMGCN, which considers background knowledge information in addition to the content of the news itself in the detection task. However, these methods inevitably ignore the social context of news, such as the (publisher-publish-news) and (user-forward-news) heterogeneous relationships, which are useful for judging fake news.

[0005] Fake news detection algorithms based on graph methods. This type of method constructs a graph model for news detection according to the characteristics of news dissemination on social platforms. The work "Beyond News Contents: The Role of Social Context for Fake News Detection" published by Shu et al. in WSDM in 2019 proposed TriFN, attempting to analyze the relationships among publishers, news, and users in social media and apply it to fake news detection. The work "Adversarial Active Learning based Heterogeneous Graph Neural Network for Fake News Detection" published by Ren et al. in ICDM in 2020 proposed AA-HGNN, which applied the method of constructing a news-oriented heterogeneous graph and then designed a graph neural network to learn news node representations. The work "Fake News Detection on News-Oriented Heterogeneous Information Networks through Hierarchical Graph Attention" published by Ren et al. in IJCNN in 2021 proposed the Hierarchical Graph Attention Network (HGAT), which achieved excellent performance in fake news discovery. The work "User Preference-aware Fake News Detection" published by Dou et al. in SIGIR in 2021 proposed UPFD, which introduced users' preference information for news in the detection task. In addition, the work "Fake News Detection with Heterogenous Deep Graph Convolutional Network" published by Kang et al. in PAKDD in 2021 proposed NDG, which successfully reduced the computational cost of the heterogeneous graph method by sampling nodes. The work "FANG: Leveraging Social Context for Fake News Detection Using Graph Representation" published by Nguyen et al. in CIKM in 2020 proposed FANG, which designed a fake news detection model that conforms to the inductive graph learning setting. However, these methods only consider the binary relationships between nodes and lack the study of high-order semantic interactions in news heterogeneous graphs, such as the relationships where two users forward the same news or a publisher publishes two news articles, thus losing key information useful for news representation learning. Summary of the Invention

[0006] Existing fake news detection graph methods usually only consider pairwise relationships between nodes, thus ignoring higher-order semantic patterns that may be more critical in news dissemination. For example, multiple users in a social network can forward the same news, a publisher can publish multiple news articles, and a user can share multiple news articles. None of these structures can be described by any binary relationship. These higher-order connection patterns are prevalent in various types of graphs (networks) and play a crucial role in graph representation learning. Therefore, once higher-order semantic structures are ignored, it may lead to the loss of key information, preventing existing heterogeneous graph methods from fully learning news representations. This situation causes the model to miss a small number of hard-to-detect fake news, so these detection methods can only achieve suboptimal detection performance.

[0007] Aiming at the problems existing in the prior art, the purpose of the present invention is to solve the suboptimal performance problem of fake news detection caused by ignoring higher-order semantic structures in most previous studies by introducing and modeling heterogeneous motif instances in social networks. A motif-based heterogeneous graph neural network fake news detection algorithm is proposed, which can accurately capture information from different heterogeneous motif instances and consider the contribution of each type of motif to fake news detection through an attention mechanism, and finally learn effective news representations for subsequent detection.

[0008] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0009] A motif-based heterogeneous graph neural network fake news detection algorithm. First, for data preprocessing, the original data from social media is constructed into a news heterogeneous graph. Second, all types of nodes are mapped to the same feature space, and corresponding instances are extracted according to each heterogeneous motif type. Then, an instance-level attention mechanism is used to aggregate all motif instances of the same type into the corresponding news nodes to capture key instance information. Next, for different types of heterogeneous motifs, a semantic-level attention mechanism is used to adaptively aggregate different news semantic embeddings. Finally, the news representation is used for the downstream fake news detection task, the prediction result of the news classification is output, and the model is continuously optimized to the optimal one. The steps are as follows:

[0010] Step (1): Data preprocessing, constructing the original data from social media into a news heterogeneous graph;

[0011] 1) Extract 3 types of nodes from the original data from social platforms: user U, news N, publisher P; and 2 types of heterogeneous binary relationships: user - forward - news U - N, publisher - publish - news P - N;

[0012] 2) Construct a news heterogeneous graph according to the nodes and relationships Among them, is a set of nodes, and ε is a set of edges. is a set of node types, is a set of edge / relationship types; the initial attribute of node v is is the type of the node, and d A is the dimension of the attributes of nodes of this type;

[0013] Step (2): Map nodes of all types to the same feature space, and extract three third-order heterogeneous motifs according to the news data, U-N-U (multiple users forwarding the same news, user-news-user), N-U-N (one user forwarding multiple news, news-user-news), and N-P-N (one publisher publishing multiple news, news-publisher-news), and extract the corresponding motif instances respectively;

[0014] 1) Use a one-layer Multi-layer Perceptron (MLP) to transform nodes of different types into the same feature space:

[0015]

[0016] where σ(·) represents a non-linear function, and ReLU(x) = max{0, x} is used as this function, and are trainable weight matrices and bias vectors, and d h is the dimension of the feature space;

[0017] 2) According to the second-order heterogeneous relationships of the news heterogeneous graph, three types of third-order heterogeneous motifs are extracted, namely U-N-U (user-news-user), N-U-N (news-user-news), and N-P-N (news-publisher-news). Then, motif instances are extracted. For node v, the set of heterogeneous motif instances of the type is that is, the set of heterogeneous motifs that contain node v and conform to the definition of m. For example, news 725 (node v)-user 28-news 892 is an instance that conforms to N-U-N. Then, the node features included in the extracted heterogeneous motif instances are concatenated. For the k-th heterogeneous motif instance, the embedding of this heterogeneous motif instance is obtained The operation of node feature concatenation is as follows:

[0018]

[0019] where (ins_1, ins_2,..., ins_n) is The node ID of the k-th heterogeneous motif instance, where the node types in the instance can be different, and CONCAT() is the concatenation operation;

[0020] Step (3): Use the attention mechanism to aggregate all heterogeneous motif instances of the same type into the corresponding news nodes to obtain the semantic representation of the nodes for the motifs;

[0021] 1) Apply a layer of MLP to calculate the attention scores of each instance related to node v, that is, the heterogeneous motif instances belonging to the same type m, and then normalize them. The calculation process is as follows:

[0022]

[0023]

[0024] Among them, is the hyperbolic tangent function, is a trainable matrix, and are trainable vectors, is the number of heterogeneous motif instances of type m related to node v;

[0025] 2) Regard the attention scores as weights, and perform a weighted sum of the instance embeddings to obtain the news representation of node v for the heterogeneous motif instances of type m. The specific calculation is as follows:

[0026]

[0027] Where is a non-linear activation function, a = 0.02, is the embedding of the k-th heterogeneous motif instance of node v;

[0028] Step (4): For different types of heterogeneous motifs, adaptively aggregate different news semantic embeddings;

[0029] 1) Evaluate the contribution of each heterogeneous motif through the attention mechanism; the semantic-level attention calculation process is as follows:

[0030]

[0031]

[0032]

[0033] Among them, is a trainable weight matrix, and are trainable weight vectors, is the number of nodes, The importance of heterogeneous motif instances of measurable type m in the fake news detection task;

[0034] 2) Use the attention scores to adaptively weight and sum all semantic representations related to node v:

[0035]

[0036] where is the number of heterogeneous motif instance types;

[0037] Step (5): Output the predicted result of the news classification and continuously optimize it to the optimal model;

[0038] 1) For fake news detection, convert it into a downstream task representation using a single-layer MLP. The downstream task representation of node v is calculated as follows:

[0039] z v = σ(h v W z + b z )

[0040] where is the trainable weight matrix, is the trainable bias vector;

[0041] 2) Obtain the predicted soft labels from the downstream task representation through the Softmax transformation. The predicted soft labels of node v are calculated as follows:

[0042]

[0043] 3) Optimize the model parameters through cross-entropy loss until the fake news detection model converges to the optimal;

[0044]

[0045] where is the set of data indices with labels in the training set, y v is the true label of node v, and there are a total of C classes; update the weight parameters through training to continuously optimize the model. After the training loss converges, the optimal algorithm model is obtained.

[0046] Compared with existing algorithms, the beneficial effects of the present invention are as follows: By introducing heterogeneous motif instances, the present invention can fuse various complex social information. These instances contain rich high-order semantic patterns that have been ignored in most previous news detection studies. Through the instance-level attention mechanism, information can be accurately captured from instances of different heterogeneous motifs, while in the semantic-level attention mechanism, the contribution of each type of motif to fake news detection is calculated, thereby learning an effective news representation and having excellent performance in fake news detection. It solves the problem of suboptimal performance in fake news detection caused by ignoring high-order semantic structures. Brief Description of the Drawings

[0047] Figure 1 is the basic framework of the present invention.

[0048] Figure 2 is the schematic diagram of the encoding and attention calculation for instances of news nodes in the present invention. Detailed Embodiment

[0049] The following further illustrates the detailed embodiments of the present invention in conjunction with the drawings and technical solutions.

[0050] To make the technical problems solved by the present invention, the technical solutions adopted, and the achieved technical effects clearer, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all the content.

[0051] It is divided into 5 steps: (1) Data preprocessing, constructing a news heterogeneous graph from the original data of social media; (2) Mapping all types of nodes to the same feature space and extracting instances according to each heterogeneous motif type respectively; (3) Using the attention mechanism to aggregate all motif instances of the same type into the corresponding news nodes to obtain the semantic representation of the nodes for the motifs; (4) Adaptively aggregating different news semantic embeddings for different types of heterogeneous motifs; (5) Outputting the prediction results of the news classification and continuously optimizing to the optimal model;

[0052] In the first step, data preprocessing, constructing a news heterogeneous graph from the original data of social media.

[0053] 1) Extract node types (user U, news N, publisher P) and relationship types (user - forward - news (U - N), publisher - publish - news (P - N)) from the original data.

[0054] 2) Construct a news heterogeneous graph according to the nodes and relationships

[0055] In the second step, map all types of nodes to the same feature space, and extract instances according to each heterogeneous motif type U-N-U, N-U-N, and N-P-N respectively.

[0056] 1) Use the formula to perform feature transformation on all types of nodes.

[0057] 2) Extract three types of heterogeneous motif instances from the heterogeneous graph, and use the formula to concatenate the node features of the k-th motif instance of node v.

[0058] In the third step, use the instance-level attention mechanism to aggregate all motif instances of the same type into the corresponding news nodes to obtain the semantic representation of the nodes for each type of motif.

[0059] 1) Calculate the attention scores for each instance (motif belonging to the same type m) related to the news node v and then normalize

[0060] 2) Use the obtained attention scores to perform weighted summation on the same-type instance embeddings, and perform a non-linear transformation using LeakyReLU.

[0061] In the fourth step, for different types of heterogeneous motifs, use the semantic-level attention mechanism to adaptively aggregate different news semantic embeddings.

[0062] 1) Evaluate the contribution of each heterogeneous motif through the attention mechanism take the average value of all news nodes and then normalize

[0063] 2) Use the semantic-level attention scores to perform weighted summation on the instance embeddings.

[0064] In the fifth step, output the prediction results of the news classification and continuously optimize to the optimal model.

[0065] 1) Use a single layer of MLP to convert it into a downstream task representation, z v = σ(h v W z + b z ).

[0066] 2) Obtain the predicted soft labels from the representation through the Softmax function.

[0067] 3) Optimize the model through cross - entropy loss until the fake news detection model of this method converges to the optimum.

[0068] Combined with the solution of the present invention, the experimental analysis is as follows:

[0069] This method uses a publicly available fake news dataset collected from social platforms to construct a news heterogeneous graph and compares it with the current main fake news detection methods to evaluate the effectiveness of this method.

[0070] (1) Introduction to the fake news detection dataset

[0071] This model conducts performance comparison tests on two publicly available social network datasets, BuzzFeed and PolitiFact, with the task of fake news detection.

[0072] The detailed information of the dataset is shown in Table 1:

[0073] Table 1 Dataset statistical information

[0074]

[0075] The BuzzFeed dataset contains 182 news, 27 publishers, and 15,257 social users. The links in BuzzFeed include N - U (news and the users who forward it) and N - P (news and the publishers who release it). Each news contains information such as the label (fake or true) of the news, the news title, the news content, the news publisher, and the history of users posting / sharing the news.

[0076] The PolitiFact dataset contains 1,056 news, 558,938 users, and 362 publishers. The links in PolitiFact include N - U (news and the users who forward it) and N - P (news and the publishers who release it).

[0077] Each news contains a unique identifier, the URL where the news is published, the news title, and the user ID of the users who share the news on Twitter. The news labels in the dataset are given. In addition, in the experiment, this method screens useful users by deleting users who forward less than two news in the dataset.

[0078] (2) Comparison experiment results of this method with other mainstream methods

[0079] The comparison of the experimental results of this method with other mainstream models in the field of fake news detection is shown in Table 2. Among them, SVM is a supervised learning method that can be used for classification, which is efficient in dealing with high-dimensional data and has applications in many fields; DW (DeepWalk) is an unsupervised graph embedding method. It constructs node sequences using random walks and projects nodes into low-dimensional embeddings using the Word2vec model, and uses a linear classifier to detect fake news; N2V (Node2vec) is similar to Deepwalk, but it applies biased random walks and has a more flexible node sequence sampling strategy, which can effectively explore different types of communities; GCN is a semi-supervised graph learning model that introduces a message passing mechanism to aggregate node neighborhood information. In the heterogeneous graph setting, all node types are ignored; GAT is a semi-supervised method in which nodes adaptively learn the representations of their neighborhoods using self-attention, and all node types are ignored in the experiment; HAN is a semi-supervised heterogeneous graph learning model. HAN aggregates the node embeddings of the meta-path-based neighborhoods through an attention mechanism and considers the importance of different meta-paths for the task; HGAT is a heterogeneous graph method for fake news detection that introduces a hierarchical two-level attention mechanism, which can aggregate the neighborhood information containing different types of nodes. The comparison experiment on BuzzFeed is shown in Table 2:

[0080] Table 2 Fake news detection results on the BuzzFeed dataset

[0081]

[0082] The comparison experiment on PolitiFact is shown in Table 3:

[0083] Table 3 Fake news detection results on the PolitiFact dataset

[0084]

[0085]

[0086] As can be seen from the results in Table 2 and Table 3, in most cases, this method has achieved the best results. For different training set ratios, this method maintains the best or competitive performance in various evaluation metrics, indicating the effectiveness and superiority of this method in the fake news detection task.

[0087] The above-described embodiments only represent the implementation manners of the present invention, but should not be construed as limiting the scope of the present invention patent. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A motif-based heterogeneous graph neural network method for fake news detection, characterized in that, The steps are as follows: Step (1): Data preprocessing, constructing a news heterogeneous graph from the original data of social media; 1) Extract 3 types of node types from the original data from social platforms: user U, news N, publisher P; and 2 types of heterogeneous binary relations: user-forward-news U-N, publisher-publish-news P-N; 2) The heterogeneous news graph constructed according to nodes and relationships , where is the set of nodes, is the set of edges, is the set of node types, is the set of edge / relationship types; the initial attribute of node is , is the type of the node, is the dimension of the attributes of nodes of this type; Step (2): Map all types of nodes to the same feature space, and extract 3 three-order heterogeneous motifs according to the news data: U-N-U, multiple users forwarding the same news, user-news-user; N-U-N, one user forwarding multiple news, news-user-news; and N-P-N, one publisher publishing multiple news, news-publisher-news; respectively extract the corresponding motif instances; 1) Use a one-layer MLP to transform different types of nodes into the same feature space: Among them, represents a non-linear function, and is used as this function, and are trainable weight matrices and bias vectors, is the dimension of the feature space; 2) According to the second-order heterogeneous relationship of the news heterogeneous graph, three types of third-order heterogeneous motifs are extracted, namely U-N-U, user-news-user, N-U-N, news-user-news, and N-P-N, news-publisher-news; then motif instances are extracted. For node , the set of heterogeneous motif instances of type is , that is, the set containing node and conforming to the heterogeneous motif defined by type ; then the node features included in the extracted heterogeneous motif instances are concatenated. For the th heterogeneous motif instance, the embedding of this heterogeneous motif instance is obtained. The operation of node feature concatenation is as follows: ​ Among them, is the node ID of the n-th heterogeneous motif instance. The node types in the instance can be different, and Step (3): Use the attention mechanism to aggregate all heterogeneous motif instances of the same type into the corresponding news nodes to obtain the semantic representation of the nodes for the motifs; Step (4): Adaptively aggregate different news semantic embeddings for different types of heterogeneous motifs; Step (5): Output the prediction results of the news classification, and continuously optimize to the optimal model.

2. The method for detecting fake news based on motif-based heterogeneous graph neural network according to claim 1, wherein Step (3) is specifically as follows: 1) Apply a single layer of MLP to compute node For each instance relevant to compute the attention scores for instances of the heterogeneous motif belonging to the same type, and then normalize them. The calculation process is as follows: Among them, is the hyperbolic tangent function, is a trainable matrix, and are trainable vectors, is the type associated with node number of motif instances; 2) Treat the attention scores as weights, perform a weighted sum on the instance embeddings, and obtain the node Regarding the type of the heterogeneous motif instance's news representation, which is calculated as follows: Among them is a non-linear activation function , is the embedding of the th heterogeneous motif instance of the node 3. The method for detecting fake news by a motif-based heterogeneous graph neural network according to claim 1, wherein, Step (4) is specifically as follows: 1) Evaluate the contribution of each heterogeneous motif through the attention mechanism; the semantic-level attention calculation process is as follows: , where is a trainable weight matrix, and are trainable weight vectors, is the number of nodes, measures the importance of heterogeneous motif instances of type in the fake news detection task; 2) Using the attention scores, adaptively weighted sum of all semantic representations related to the node : wherein is the number of heterogeneous motif instance types.

4. A motif-based heterogeneous graph neural network-based fake news detection method according to claim 1, characterized in that, Step (5) is specifically as follows: 1) For fake news detection, convert it into a downstream task representation using a single layer of MLP, and the downstream task representation of node is calculated as follows: Among them, is a trainable weight matrix, is a trainable bias vector; 2) Through the Softmax transformation, obtain the predicted soft labels from the downstream task representations. The predicted soft label of node is calculated as follows: , 3) Train and optimize the model parameters through the cross-entropy loss until the fake news detection model converges to the optimal; Among them, is the set of data index with labels in the training set, is the true label of node , and there are categories in total; by training and updating the weight parameters, the continuous optimization of the model can be achieved. After the training loss converges, the optimal algorithm model can be obtained.

Citation Information

Patent Citations

  • Joint false news detection method based on mode information and fact information

    CN113849599A

  • False news identification method based on heterogeneous graph contrast learning

    CN114020928A