A False News Detection Method for Graph Neural Networks Based on Contrastive Learning
By combining GCN and GAT, introducing residual connections and contrast learning, and constructing propagation and diffusion maps, the neglected problem of propagation structure and diffusion structure in fake news detection is solved, and the accuracy and robustness of the detection are improved.
Patent Information
- Application Number
- CN202311364692.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-10-20
AI Technical Summary
The existing fake news detection methods are insufficient in terms of efficiency and accuracy, especially ignoring the dissemination structure and diffusion structure of fake news in social networks, and the loss function is flawed in category imbalance and text context understanding.
A comparison learning method is used to combine GCN and GAT to construct propagation graphs and diffusion graphs, introduce residual connections, extract features through graph neural networks, and jointly train using comparison loss and classification loss functions to capture event invariant features.
Improves the robustness and accuracy of fake news detection, can better mine and utilize local information in the graph, capture longer-distance dependencies, and improves category imbalance problems and text context understanding.
Smart Images

Figure CN117195080B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fake news detection, and particularly to a fake news detection method based on contrastive learning of graph neural networks. Background Art
[0002] The popularity of social media has accelerated the spread of fake news, which has serious negative impacts on society, politics, and the economy. These false information are widely spread on social media, online platforms, and news media, which may lead to the public's misunderstanding of facts, the spread of rumors, and the guidance of false public opinions. The importance of researching fake news detection technology is obvious. It helps to improve network information security, protect the public's right to know and freedom of speech. Timely discovery and response to fake news contribute to maintaining social stability and preventing potential social risks. Therefore, effectively detecting and identifying fake news has become an urgent need.
[0003] So far, many methods have achieved impressive results in fake news detection. For example, Long Short-Term Memory (LSTM) can learn sequence features from the spread of fake news over time. However, these methods have great limitations in terms of efficiency because the time structure features only focus on the sequential spread of fake news and ignore the impact of the spread of fake news. Since the spread patterns of real news and fake news in social networks are different, the structure of the spread of fake news also reflects some spreading behaviors of rumors, which is conducive to classifying fake news. For example, Bi-GCN uses a top-down graph to represent the dissemination information of news, a bottom-up graph to represent the diffusion information of news, and GCN to fuse the node information in the graph to obtain node representations, achieving good classification results. However, GCN uses a fixed weight matrix to aggregate neighbor nodes in each layer, ignoring the differences between different nodes. Moreover, GCN only considers the first-order neighbor nodes of nodes in each layer and cannot directly obtain more distant association information. This deficiency inspires us that the Graph Attention Network (GAT) can also be applied to the model. Because GAT can adaptively learn the importance of different neighbor nodes by introducing an attention mechanism, and by weighting nodes at different distances, the model can more comprehensively perceive the information of the entire graph, thereby better capturing the relationships between nodes. However, in the deep neural network model that fuses GCN and GAT, with the increase in the number of layers, the problem of gradient disappearance or gradient explosion may become more serious, leading to difficulties in the model training process. This problem inspires us that by introducing residual connections in the network, it can help information or gradients to directly propagate across layers, effectively alleviating this problem.
[0004] In terms of the loss function, the negative log-likelihood loss is widely used in classification problems, but it also has some potential drawbacks. For example, it is sensitive to class imbalance. In the fake news detection task, the ratio of real fake news to real real news may be imbalanced. In this case, these loss functions may be biased towards the class with more samples, thus affecting the performance of the model on the minority class. In addition, in fake news detection, the context and context of the text may be crucial for judging whether it is fake news or not, but these loss functions usually only consider the difference between the model's output and the true label, and do not consider that the context information of the text cannot directly capture this information. For these limitations and drawbacks, contrastive learning can introduce positive and negative sample pairs between samples to improve the class imbalance problem. It can help the model better distinguish between real news and fake news, especially in the case of data imbalance. At the same time, because they usually need to compare the similarity between two or more samples, this can help improve the model's understanding and modeling of text context. Summary of the Invention
[0005] The object of the present invention is to provide a method for detecting fake news based on a graph neural network with contrastive learning, aiming at the deficiencies of the prior art. The method combines GCN and GAT and introduces residual connections to better mine and utilize local information in the graph while capturing longer-distance dependencies, making the model more robust and effective. This method not only considers the propagation structure information of rumors but also the diffusion structure information of rumors, greatly improving the robustness of the model and the accuracy of fake news detection, and having good application prospects.
[0006] The specific technical solution for implementing the present invention is: a method for detecting fake news based on a graph neural network with contrastive learning, characterized in that this method combines GCN and GAT and introduces residual connections to better mine and utilize local information in the graph while capturing longer-distance dependencies, specifically including the following steps:
[0007] S1. Construct the propagation graph and diffusion graph of the news according to the propagation relationship of the news posts and perform data augmentation;
[0008] S2. Obtain the top-down propagation features and bottom-up diffusion features through the graph neural network;
[0009] S3. Construct positive and negative sample pairs to make samples with the same label gather together and samples with different labels be pushed away, and use the loss function of contrastive learning for the training of the graph neural network model to achieve similarity learning;
[0010] S4. Aggregate the propagation features and diffusion features, then make connections, and then predict event labels through the connection layer and the Softmax layer. Use the classification loss function to train the graph neural network model. The total loss function includes the loss function of contrastive learning and the loss function of classification. Classify the news using the trained graph neural network model to distinguish between real news and fake news.
[0011] The specific steps of step S1 include: constructing a news event c according to the forwarding and response relationships of news posts i 's propagation structure <V i , E i >. Let and X be its corresponding adjacency matrix and the feature matrix of c based on the news post propagation tree i . A only contains edges from the upper node to the lower node. It is shown that in each training stage, p percentage of the edges are deleted through the equation A' = A - A drop . In this way, the problem of penalized overfitting can be avoided. Based on A' and X, the present invention can create the propagation graph and diffusion graph of each news post. The adjacency matrix of the propagation graph is A TD , A TD = A', and the adjacency matrix of the diffusion graph is A BU , A BU = A' T . The propagation graph and the diffusion graph adopt the same feature matrix X;
[0012] The goal of step S2 is to obtain the top-down propagation features and the bottom-up diffusion features, which specifically include:
[0013] S2-1. The graph neural network layer used in the present invention consists of two parts: P-GNN and D-GNN, namely the Propagated Graph Neural Network and the Dispersive Graph Neural Network. Their network architectures are the same and share weights. Both P-GNN and D-GNN are composed of a 3-layer graph neural network. The first two layers are mainly graph convolutional neural network GCN, and the third layer is mainly graph attention network GAT. A residual structure is added between each layer of the network.
[0014] S2-2. Graph convolutional neural network GCN: Since GCN has multiple information propagation functions M, the message propagation function defined in its first-order approximation is shown in the following formula (a):
[0015]
[0016] In the above equation, is the normalized adjacency matrix, where (i.e., adding self-connection), represents the degree of the first node; σ(·) is an activation function.
[0017] In the present invention, by substituting A TD and X of the news dissemination graph data into formula (a), the features extracted by the first-layer graph convolutional neural network GCN of P-GNN can be obtained, as shown in the following formula (b):
[0018]
[0019] Then, by substituting and X into the generalized formula of the residual structure, the complete feature representation of the first-layer graph convolutional neural network GCN of P-GNN can be obtained, as shown in the following formula (c):
[0020]
[0021] Similarly, the features extracted by the second and third-layer graph convolutional neural networks GCN of P-GNN are as shown in the following formulas (d) and (e):
[0022]
[0023]
[0024] S2-3. Use the outputs of the first 2-layer graph convolutional neural networks GCN as the inputs of the third-layer graph attention network layer. The GAT network in the third-layer graph attention network layer will consider the attention coefficients of all nodes adjacent to node i. Its key formula is as shown in the following formula (f):
[0025]
[0026] By Through formula (f), different weights can be assigned to the neighbor nodes of each node, the attention coefficients can be calculated, and then combined with the generalized formula of the residual structure, the features extracted by the final graph neural network of P-GNN can be obtained, specifically as shown in the following formula (g):
[0027]
[0028] Among them, and represent the hidden features of the propagation graph in the three-layer graph neural network of P-GNN; and is the parameter matrix of P-GNN. The present invention uses the ReLU function as the activation function: σ(·); DropEdge is applied to the graph convolutional neural network GCN layer and the graph attention network layer to avoid overfitting.
[0029] S2-4. Similar to the formulas (b) to (g), the present invention uses the same method as the equalities (b) to (g) to input A BU and X of the diffusion map data of the news into the D-GNN to calculate the bottom-up hidden features of the diffusion map: and
[0030] The objective of step S3 is to construct positive and negative sample pairs so that samples with the same label are clustered together and samples with different labels are pushed apart. The contrast loss function is used for training to achieve similarity learning:
[0031] S3-1. During the training process, the data volume of the data K in the same batch is N, and a data k, k ∈ K ≡ {1…N} is selected.
[0032] S3-2. This k is called the anchor. Select the data with the same label as k (the same category) in the data K of the same batch, calculate the cosine similarity with k, and select the most similar data as p, which is used as the positive sample of k. Select the data with a different label from k (the different category) in the data K of the same batch, calculate the cosine similarity with k, and select the data with the lowest similarity as a, which is used as the negative sample of k.
[0033] S3-3. Convert these samples into vector representations in the high-dimensional space through the graph neural network (see step S2 in detail), and use the Euclidean distance between vectors to measure the similarity of samples, so that the distance between the anchor and the positive example is as small as possible, while the distance between the anchor and the negative sample is as large as possible. The contrast loss function used is as follows in formula (h):
[0034]
[0035] where, m k is the anchor, if m a and m k have different labels, then they are used as negative samples, and the negative sample set can be expressed as: A(k) ≡ {a ∈ K: m a ≠ m k}; if m p and m k have the same label, then they are used as positive samples, and the positive sample set is expressed as: P(k) ≡ {p ∈ K: m p ≠ m k; label is a tag indicating whether two samples belong to the same category. If they belong to the same category, that is, when the input is a positive sample pair, label is 0; if they belong to different categories, that is, when the input is a negative sample pair, then it is 1; m represents the minimum distance between samples with the same label (positive samples) and samples with different labels (negative samples), that is, if the distance between two samples of the same category is less than m, they are considered "similar", otherwise they are considered "dissimilar"; d is used to calculate the Euclidean distance between samples, that is, the distance between two samples in the feature space.
[0036] S3-4. The calculation of the loss function is divided into two parts, corresponding to samples of the same category and samples of different categories respectively. For samples of the same category (label is 0), calculate the square of the Euclidean distance, and the goal is to make the distances of these samples as close to 0 as possible. For samples of different categories (label is 1), calculate the margin loss, that is, calculate the difference between the minimum distance threshold between negative sample pairs and their actual Euclidean distance (m - d). When the difference is positive, that is, the negative sample pairs are too close, the loss is greater than 0; when the difference is negative, that is, the negative sample pairs are far enough away from each other, the loss is set to 0. Furthermore, by reducing the loss, the goal of clustering positive sample pairs and separating negative sample pairs can be achieved. Finally, add the two parts of the loss and then take the average to obtain the final contrastive loss.
[0037] The goal of step S4 is to aggregate the propagated features and diffused features, then make a connection, and then predict the event label through the connection layer and the Softmax layer, and use the classification loss function for training. The total loss function includes: contrastive learning and classification loss; step S4 specifically includes:
[0038] S4-1. Aggregate the top-down propagated feature H TD obtained from the graph neural network of CGAR and the bottom-up diffused feature H BU respectively, and use the average pooling operator to aggregate the information from these two groups of nodes, as shown in the following formula (i):
[0039]
[0040] S4-2. Connect the aggregated propagated representation and diffused representation, and merge the information into the following formula (j):
[0041] M = concat(M TD , M BU ) (j).
[0042] S4-3. The label of the event is calculated by the following formula (k) through several fully connected layers and a Softmax layer:
[0043]
[0044] Among them, is the probability vector of all classes for predicting event labels.
[0045] S4-4. By minimizing the cross-entropy between the prediction and the actual distribution Y, for all events C, all parameters in the contrast graph attention residual (CGAR) model are trained. In the loss function, L2 regularization is also applied to all model parameters, and the classification loss function is shown in the following formula (m):
[0046]
[0047] Among them, λ is the regularization coefficient, and p is the CGAR model parameter.
[0048] It can be obtained from the above that the total loss function is the loss function of contrast learning plus the loss function of classification, that is, calculated by the following formula (n):
[0049] L total = L Contrastive + L Classification (n).
[0050] Compared with the prior art, the present invention has better robustness and greatly improves the accuracy of fake news detection. The present invention not only considers the propagation structure information of rumors, but also considers the diffusion structure information of rumors, captures event invariant features using contrast learning, performs joint training using contrast loss and classification loss, combines GCN and GAT, and introduces residual connections, which can help the model better mine and utilize local information in the graph while capturing longer-distance dependencies, making the model more robust and effective, and having good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is the flowchart of the present invention;
[0052] Figure 2 is the schematic diagram of the framework of the present invention;
[0053] Figure 3 is the model performance graph in the parameter exploration stage;
[0054] Figure 4 is the experimental result comparison graph for early fake news detection;
[0055] Figure 5 is the example graph of the experimental results for gradually excluding components or functions of the model. DETAILED DESCRIPTION OF THE INVENTION
[0056] In conjunction with the following specific drawings, the present invention will be further described in detail. The processes, conditions, experimental methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge and well-known common sense in the art. For those of ordinary skill in the art, without creative labor, other drawings can be obtained based on these drawings, and other implementation manners can also be obtained.
[0057] Example 1
[0058] Refer to Figures 1 to 2 , the present invention performs false news detection of graph neural networks based on contrastive learning according to the following steps:
[0059] S1. Construct the propagation graph and diffusion graph of the news according to the propagation relationship of the news posts and perform data augmentation;
[0060] S2. Obtain the top-down propagation features and bottom-up diffusion features through the graph neural network;
[0061] S3. Construct positive and negative sample pairs to cluster samples with the same label together and push samples with different labels away, and use the loss function L of contrastive learning Contrastive for training the graph neural network model to achieve similarity learning;
[0062] S4. Aggregate the propagation features and diffusion features and perform connection, predict the event label through the connection layer and the Softmax layer, and use the classification loss function L Classification to train the graph neural network model, and classify the news with the trained graph neural network model to distinguish real news and false news.
[0063] In step S1: According to the forwarding and response relationships of the news posts, construct a propagation structure <V i , E i > of a news event c i . Let and X be the corresponding adjacency matrix and the feature matrix of c based on the news post propagation tree i respectively. A only contains the edges from the upper node to the lower node. It is shown that in each training stage, p percentage of the edges are deleted through the equation A' = A - A drop . In this way, the problem of penalized overfitting can be avoided. Based on A' and X, the present invention can create the propagation graph and diffusion graph of each news post. The adjacency matrix of the propagation graph is A TD , A TD = A', and the adjacency matrix of the diffusion graph is A BU , A BU = A' T . The propagation graph and the diffusion graph adopt the same feature matrix X;
[0064] Step S2: Obtain top-down propagation features and bottom-up diffusion features, which specifically include:
[0065] S2-1. The graph neural network layer of the present invention consists of two parts: P-GNN and D-GNN, namely the Propagated Graph Neural Network and the Dispersive Graph Neural Network. The network architectures of both are the same and share weights. Both P-GNN and D-GNN are composed of a 3-layer graph neural network. The first two layers are mainly graph convolutional neural network GCN, and the third layer is mainly graph attention network GAT. A residual structure is added between each layer of the network.
[0066] S2-2. Graph convolutional neural network GCN: Since there are multiple information propagation functions M in GCN, the message propagation function defined in its first-order approximation is shown in the following formula (a):
[0067]
[0068] In the above equation, is the normalized adjacency matrix, where (that is, adding self-connections), represents the degree of the first node; σ(·) is an activation function; in the present invention, by substituting A TD and X of the news propagation graph data into formula (1), the features extracted by the first-layer graph convolutional neural network GCN of P-GNN can be obtained, as shown in the following formula (b):
[0069]
[0070] Then, by substituting and X into the generalized formula of the residual structure, the complete feature representation of the first-layer graph convolutional neural network GCN of P-GNN can be obtained, as shown in the following formula (c):
[0071]
[0072] Similarly, the features extracted by the second and third-layer graph convolutional neural networks GCN of P-GNN are shown in the following formulas (d) and (e):
[0073]
[0074]
[0075] S2-3. Use the output of the first two layers of the graph convolutional neural network (GCN) as the input to the third layer of the graph attention network layer. The GAT network in the third layer of the graph attention network layer will consider the attention coefficients of all nodes adjacent to node i for node i. Its key formula is shown in the following formula (f):
[0076]
[0077] Use Through formula (f), different weights can be assigned to the neighbor nodes of each node to calculate the attention coefficients. Then, combined with the general formula of the residual structure, the features extracted by the final graph neural network of P-GNN can be obtained, as shown in the following formula (g):
[0078]
[0079] Among them, And represent the hidden features of the propagation graph in the three-layer graph neural network of P-GNN; And are the parameter matrices of P-GNN. Here, the ReLU function is used as the activation function: σ(·); DropEdge is applied to the GCN layer and the graph attention network layer of the graph convolutional neural network to avoid overfitting.
[0080] S2-4. Similar to equations (b) to (g), the present invention uses the same method as equations (b) to (g) to input the A BU and X of the diffusion graph data of the news into D-GNN to calculate the bottom-up hidden features of the diffusion graph: And
[0081] S3: Construct positive and negative sample pairs to gather samples with the same label together and push samples with different labels away. Use the contrast loss function for training to achieve similarity learning, specifically including:
[0082] S3-1. During the training process, the data volume of the data K in the same batch is N, and a data k, k ∈ K ≡ {1…N} is selected.
[0083] S3-2. Denote this k as the anchor. Among the data K in the same batch, select the data with the same label as k (the same category), calculate the cosine similarity with k, and select the most similar data as p, which serves as the positive sample of k. Among the data K in the same batch, select the data with a different label from k (different categories), calculate the cosine similarity with k, and select the data with the lowest similarity as a, which serves as the negative sample of k.
[0084] S3-3. Convert these samples into vector representations in a high-dimensional space through a graph neural network (see step S2 in detail). Use the Euclidean distance between vectors to measure the similarity of samples, making the distance between the anchor and the positive example as small as possible, while the distance between the anchor and the negative sample as large as possible. The contrastive loss function used is shown in the following formula (h):
[0085]
[0086] where m k is the anchor, if m a and m k have different labels, then they serve as negative samples, and the negative sample set can be expressed as: A(k)≡{a∈K:m a ≠m k}; if m p and m k have the same label, then they serve as positive samples, and the positive sample set is expressed as: P(k)≡{p∈K:m p ≠m k}; label is a label indicating whether two samples belong to the same category. If they belong to the same category, that is, when the input is a positive sample pair, label is 0; if they belong to different categories, that is, when the input is a negative sample pair, then it is 1; m represents the minimum distance between samples with the same label (positive samples) and samples with different labels (negative samples). In other words, if the distance between two samples of the same category is less than m, they are considered "similar", otherwise they are considered "dissimilar"; d is used to calculate the Euclidean distance between samples, that is, the distance between two samples in the feature space.
[0087] S3-4. The calculation of the loss function is divided into two parts, corresponding to the same-class samples and different-class samples respectively. For the same-class samples (label = 0), calculate the square of the Euclidean distance, and the goal is to make the distances of these samples as close to 0 as possible. For different-class samples (label = 1), calculate the boundary loss, that is, calculate the difference between the minimum distance threshold between negative sample pairs and their actual Euclidean distance (m - d). When the difference is positive, that is, the negative sample pairs are too close, the loss is greater than 0; when the difference is negative, that is, the negative sample pairs are far enough away from each other, the loss is set to 0. Furthermore, by reducing the loss, the goal of clustering positive sample pairs and separating negative sample pairs can be achieved. Finally, add the two parts of the loss and then take the average to obtain the final contrastive loss.
[0088] S4: Aggregate the propagation features and diffusion features, and then connect them. Then, through the connection layer and the Softmax layer, predict the event labels, and use the classification loss function for training. The total loss function includes: contrastive learning and classification loss. The specific steps of S4 are as follows:
[0089] S4-1. Aggregate the top-down propagation features H TD and the bottom-up diffusion features H BU obtained from the graph neural network of CGAR respectively. Use the average pooling operator to aggregate the information from these two groups of nodes, as shown in the following formula (i):
[0090]
[0091] S4-2. Connect the aggregated propagation representation and diffusion representation, and merge the information as shown in the following formula (j):
[0092] M = concat(M TD , M BU ) (j).
[0093] S4-3. The label of the event is calculated through several fully connected layers and a Softmax layer as shown in the following formula (k):
[0094]
[0095] where is the probability vector of all classes used to predict the event label.
[0096] S4-4. By minimizing the cross-entropy between the prediction and the actual distribution Y, for all events C, train all the parameters in the contrastive graph attention residual (CGAR) model. In the loss function, L2 regularization is also applied to all model parameters. The classification loss function is as shown in the following formula (m):
[0097]
[0098] Among them, λ is the regularization coefficient and p is the CGAR model parameter.
[0099] From the above, the total loss function is the loss function of contrastive learning plus the loss function of classification, as shown in the following formula (n):
[0100] L total = L Contrastive + L Classification (n).
[0101] To verify the generality of the present invention, eight different typical networks are selected: LIWC, VGG-19, att-RNN, SAFE, Bi-GCN, GNN-CL, GCNFN, UPFD, for comparative experiments. To conduct a fair comparison, the present invention randomly partitions the dataset, which can ensure that the data in each subset is representative, thus avoiding imbalance caused by data sorting or some bias. For the two datasets of PolitiFact and GossipCop, the present invention evaluates the accuracy (Acc), precision (Pre), recall (Rec), and F1. By using stochastic gradient to update the parameters of the CGAR model, the model is optimized by the Adam algorithm. For the PolitiFact dataset, the graph embedding size (128), the hidden layer size of the node is (32), the batch size (128), the optimizer (Adam), and the L2 regularization weight (0.01) are used, and the dropout rate in DropEdge is 0.2. For the GossipCop dataset, the graph embedding size (128), the hidden layer size of the node is 128, the batch size (128), the optimizer (Adam), and the L2 regularization weight (0.002) are used, and the dropout rate in DropEdge is 0.5. The training process iterates for 100 epochs. The training-validation-test split is (60%-20%-20%). The experimental results are averaged over five different runs.
[0102] The performance of the present invention and the method for comparison therewith on the PolitiFact and GossipCop datasets is shown in Table 1 below:
[0103] Table 1 False news detection results on the Politifact and Gossipcop datasets
[0104]
[0105]
[0106] Table 1 above shows the performance of the present invention and the methods compared therewith on the PolitiFact and GossipCop datasets. First, among the benchmark algorithms, we observe that the performance of deep learning methods is significantly better than that of methods using handcrafted features. This is not surprising because deep learning methods can learn high-level representations of news to capture effective features. This demonstrates the importance and necessity of studying deep learning for fake news detection. Second, the accuracy, precision, recall, and F1-score of the proposed method CGAR of the present invention exceed those of methods such as LIWC, VGG-19, GNN-CL, GCNFN, UPFD, and Bi-GCN. This result further confirms the effectiveness and superiority of the CGAR method. In addition, the experiment also conducts a deep analysis of the CGAR method. Through the results, it is found that the advantages of the CGAR method mainly come from the combination of its contrastive learning and graph attention residual network. Contrastive learning enables the model to better understand and distinguish the differences between real news and fake news, while the graph attention residual network enables the model to better capture and utilize the relationships between news. The combination of these two factors makes the CGAR method show significant advantages in the fake news detection task.
[0107] Refer to Figure 3 (a) to Figure 3 (b), the nhid parameter represents the number of neurons in the hidden layer, that is, the size of the hidden layer. In a neural network, the number of neurons in the hidden layer determines the complexity and learning ability of the model. The larger the value of the nhid parameter, the stronger the complexity and learning ability of the model, but at the same time, it may also increase the risk of overfitting. In the experimental results (as Figure 3 shown), the Politifact dataset performs best when nhid is 32. When nhid is 32, it performs best. When nhid is 64, 128, or 256, the overfitting becomes more and more serious and the performance gradually decreases. When nhid is 16, it is underfitting. For the Gossipcop dataset, it performs best when nhid is 128. When nhid is 128, it performs best. When nhid is 256, it is overfitting. When nhid is 64, 32, or 16, the underfitting becomes more and more serious and the performance gradually decreases. This may be because the characteristics of these two datasets are different and models with different complexities are required to best capture their features. The Politifact dataset has a small amount of data and relatively few features, so a simpler model (i.e., a smaller nhid) can capture these features well and is not easily overfitted. The Gossipcop dataset has a larger amount of data and its internal structure or features are more complex, so a more complex model (i.e., a larger nhid) is required to capture these features.
[0108] The main goal of DropEdge is to alleviate the overfitting problem by randomly dropping edges in the graph during training, thereby improving the generalization ability of the model and enhancing the robustness of the model. This method can be regarded as an extension of the Dropout method for graph neural networks. The main parameter of DropEdge is the dropping probability, usually between 0 and 1.
[0109] See Figure 3 (c) to Figure 3 (d). In the experimental results, the Politifact dataset performs best when DropEdge is 0.2, while the Gossipcop dataset performs best when DropEdge is 0.5. It may be because there are a large number of redundant or unimportant edges in the graph structure of the gossipcop dataset, and these edges may interfere with the model's learning of the true relationships in the data. By increasing the "DropEdge" parameter, a part of the edges can be randomly deleted during training, thereby reducing the impact of these redundant or unimportant edges on the model and improving the performance of the model. For the larger dataset Gossipcop, the model can usually have a larger capacity (more parameters or a more complex structure) because there is enough data to learn these parameters. Therefore, using a higher DropEdge rate can help regularize larger and more complex models. A higher DropEdge rate forces the model not to rely on any specific neurons, thereby improving the robustness. This is beneficial for large models and large datasets because it encourages the model to learn multiple representations from the data, which can help the model better generalize to unseen data. For the small dataset Politifact, it may be because although the Politifact dataset has a small number, its graph structure is complex and requires more node information to capture the relationships in the data. A smaller "DropEdge" parameter means that more nodes are retained during training, enabling the model to learn more complex features. However, if the "DropEdge" parameter is too large, it may cause the model to lose important node information and fail to capture the complex relationships in the data, resulting in a performance decline.
[0110] Weight Decay: Weight Decay is a technique that penalizes the model parameters, usually used to prevent the model's weights from becoming too large. In the model's loss function, in addition to the original loss term, the sum of the squares of the weights multiplied by a coefficient (i.e., the Weight Decay coefficient) is added as a penalty term. This can prevent the model's weights from becoming too large, thereby preventing overfitting. However, it may also lead to underfitting. The main parameter of Weight Decay is the penalty coefficient. It and "DropEdge" are both regularization techniques, but the roles they play and the ways they handle problems in model training are different. The complementarity of these two techniques lies in that they are both designed to improve the generalization ability of the model, but their working methods are completely different. DropEdge achieves this goal by operating on the data, while weight_decay achieves this goal by adjusting the complexity of the model. Therefore, they can be used simultaneously to improve the generalization ability of the model from different perspectives.
[0111] See Figure 3 (e)~ Figure 3 (f), in the experimental results, the Politifact dataset performs best when weight_decay is 0.01, while the Gossipcop dataset performs best when weight_decay is 0.002. This may be because the Politifact dataset has a certain degree of noise and requires stronger regularization to prevent overfitting. Moreover, Politifact is a small dataset, which means that the model has a higher risk of overfitting. When the model's weights become larger, it becomes more dependent on specific patterns and noise in the training data. By increasing weight_decay, the model will be penalized and tend to learn smaller weights, thereby reducing the risk of overfitting. And a higher weight_decay will cause the model weights to converge to zero, resulting in a simpler and sparser model. For a small dataset, a simplified model is usually easier to generalize. For the larger dataset Gossipcop, the smaller the "weight_decay" parameter, the better the performance. This may be because the overall complexity of this dataset is relatively high and a more complex model is needed to capture the relationships in the data.
[0112] Early fake news detection is a technique for identifying and preventing the spread of false information as early as possible. The main goal of this technique is to accurately identify a rumor before it starts to spread and has a potential impact on society. Therefore, it is very necessary to conduct early fake news detection experiments. The present invention sets a series of detection deadlines and only uses the posts published before the deadlines to evaluate the accuracy of the proposed method and the baseline method.
[0113] SeeFigure 4 , the performance of the CGAR method of the present invention relative to other false news detection models based on graph neural networks, such as GNN-CL, GCNFN, UPFD, and Bi-GCN, at different deadlines on the Gossipcop and Politifact datasets. It can be seen from the figure that the proposed CGAR method achieves high accuracy in the early stage after the source initial broadcast. In addition, the overall performance of CGAR in terms of ACC, PRE, REC, and F1 at each deadline is significantly better than that of other models, indicating that structural features are not only beneficial for long-term false news detection but also contribute to the early detection of false news.
[0114] CGAR-C represents the variant model of CGAR after removing the contrastive learning module, and CGAR-gar represents the variant model of CGAR after removing the graph attention residual module respectively.
[0115] Refer to Figure 4 (a) to Figure 4 (h), it can be seen that in the early detection of false news, as time goes by, the performance of the CGAR model, CGAR-c model, and CGAR-gar model all gets better and better. At each time point in the early detection, the performance of the CGAR model is better than that of the CGAR-c model and CGAR-gar model. At the earlier time, the performance gap between the CGAR-c model and CGAR-gar model and the CGAR model is relatively large, and as time goes by, the gap becomes smaller and smaller. Through the above experimental data, we can see that both the contrastive learning module and the graph attention residual module of the CGAR model play important roles in improving the performance of the entire model. In the false news detection task, the contrastive learning module can effectively improve the discriminative ability of the model by learning how to distinguish real news from false news. When the contrastive learning module is removed, the model may lose this discrimination ability, resulting in a performance decline. The graph attention residual module can dynamically assign weights to the importance of nodes, enabling the model to pay more attention to important nodes. Moreover, through the residual connection, it can help the model better learn the mapping relationship between the input and output, avoiding the problems of gradient disappearance and explosion, and enabling the model to learn more deeply. When the graph attention residual module is removed, the model may lose this ability to capture graph structure information and deep learning, resulting in a performance decline.
[0116] Refer to Figure 5 (a) to Figure 5(b), it can be seen that in the long-term fake news detection, for the Politifact dataset, compared with the CGAR-c model, the CGAR model has improved the performance of Acc, Pre, Rec, and F1 by 3.2%, 0.3%, 7.3%, and 3.8% respectively; compared with the CGAR-gar model, the performance of Acc, Pre, Rec, and F1 has increased by 1.6%, 0.1%, 4.8%, and 2.2% respectively. For the Gossipcop dataset, compared with the CGAR-c model, the CGAR model has improved the performance of Acc, Pre, Rec, and F1 by 2.7%, 0.1%, 5.6%, and 2.9% respectively; compared with the CGAR-gar model, the performance of Acc, Pre, Rec, and F1 has increased by 2.4%, 0.7%, 4.4%, and 2.5% respectively.
[0117] The above is only a detailed description of the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.
Claims
1. A method for detecting false news in graph neural networks based on contrastive learning, characterized in that, The method specifically includes the following steps: S1. Construct a propagation graph and a diffusion graph of the news according to the propagation relationship of the news posts and perform data augmentation; S2. Obtain top-down propagation features and bottom-up diffusion features through a graph neural network. The graph neural network consists of P-GNN and D-GNN, and their network architectures are the same and share weights. Both P-GNN and D-GNN are composed of graph neural networks with three-layer structures. The main bodies of the first two layers are graph convolutional neural networks GCN, and the main body of the third layer is a graph attention network GAT. Moreover, a residual structure is added between each layer of the network; S3. Construct positive and negative sample pairs to cluster samples with the same label together and push samples with different labels apart, and use the loss function \(L\) of contrastive learning Contrastive for training the graph neural network model to achieve similarity learning; S4. Aggregate the propagation features and diffusion features and perform connection, predict the event labels through the connection layer and the Softmax layer, and use the classification loss function L Classification Train the graph neural network model, and classify news using the trained graph neural network model to distinguish between real news and fake news.
2. The method for detecting fake news in a graph neural network based on contrastive learning according to claim 1, characterized in that, The specific steps of S1 include: S1-1: Construct a propagation structure <V i , E i > of a news event c according to the forwarding and response relationships of news posts. Let A and X be the corresponding adjacency matrix and the eigenmatrix of c based on the news post propagation tree, respectively, where A is the eigenmatrix containing only the edges from the upper nodes to the lower nodes; i be the dimension of the adjacency matrix Ai; V i be the node set of the graph; E i be the edge set of the graph. i ,E i >, set and X be the corresponding adjacency matrix and the eigenmatrix of c based on the news post propagation tree, respectively, where A is the eigenmatrix containing only the edges from the upper nodes to the lower nodes; i of c; be the dimension of the adjacency matrix Ai; V i be the node set of the graph; E i be the edge set of the graph; S1-2: In each training stage, use DropEdge for data augmentation as follows: Delete p percentage of edges through the A′ = A - A drop equation, where A' is the adjacency matrix after calculating DropEdge; p is the deletion percentage; A drop is the matrix constructed from newly sampled edges from the original edge set; S1-3: Create the propagation graph and diffusion graph for each news post based on A' and X, where the adjacency matrix of the propagation graph is A TD , and A TD = A'; the adjacency matrix of the diffusion graph is A BU , and A BU = A 'T ; the propagation graph and diffusion graph adopt the same feature matrix X.
3. The method for detecting fake news in a graph neural network based on contrastive learning according to claim 1, wherein The specific steps of S2 include: S2-1. The graph convolutional neural network GCN has multiple information propagation functions M. The message propagation function defined in its first-order approximation is defined by the following formula (a): H k = M(A, H k-1 ; W k-1 ) = σ(AH k-1 W k-1 ) (a); where M is the information propagation function, A is the adjacency matrix, H k-1 is the hidden feature matrix, is the trainable parameter matrix, is the normalized adjacency matrix, A = A + I N is the added self-connection; D ii = Σ j A ij is the degree of the first node; σ(·) is the activation function; Substitute A of the news dissemination graph data TD and X into the above formula (a) to obtain the features extracted by the first-layer graph convolutional neural network GCN of P-GNN shown in the following formula (b): Among them, A TD is the standardized adjacency matrix; X is the feature matrix; is the trainable parameter matrix; σ(·) is the activation function; By substituting and X into the generalized formula of the residual structure, the complete feature representation of the first-layer graph convolutional neural network GCN of P-GNN shown in the following formula (c) is obtained: Similarly, the features extracted by the second-layer graph convolutional neural network GCN of P-GNN shown in the following formulas (d) and (e) can be obtained: Among them, A TD is a standardized adjacency matrix; is the complete feature representation of the first-layer graph convolutional neural network GCN of P-GNN; W1 TD is a trainable parameter matrix; σ(·) is an activation function; is the feature representation extracted by the second-layer graph convolutional neural network GCN of P-GNN; S2-2. Take the outputs of the first 2 layers of the graph convolutional neural network GCN as the inputs of the third-layer graph attention network layer. The GAT network in the third-layer graph attention network layer will consider the attention coefficients of all nodes adjacent to node i, and is defined by the following formula (f): where, σ(·) is the activation function; N i is the domain composed of all nodes adjacent to node i; is the feature vector of node j; W is a weight matrix of size F′×F; is the attention relationship between two nodes; e ij is the attention coefficient to represent the influence of node i on node j; The complete feature representation parameters of the second-layer graph convolutional neural network GCN of P-GNN According to equation (f), different weights are assigned to the neighbor nodes of each node to calculate the attention coefficients, and then combined with the general formula of the residual structure, the features extracted by the final graph neural network of P-GNN shown in the following equation (g) are obtained: Among them, and are the hidden features of the propagation graph in the three-layer graph neural network of P-GNN; and are the parameter matrices of P-GNN, and the ReLU function is used as the activation function σ(·); DropEdge is applied to the graph convolutional neural network GCN layer and the graph attention network layer; is the feature extracted by the second-layer graph convolutional neural network GCN of P-GNN; is the feature extracted by the third-layer graph attention network GAT of P-GNN; S2-3: Using the same method as in formulas (b) to (g), input A of the diffusion map data of the news BU and X into the D-GNN, and calculate the bottom-up hidden features of the diffusion map: and 4. The method for detecting fake news in a graph neural network based on contrastive learning according to claim 1, wherein, The specific steps of S3 include: S3-1. During the training process, the data volume of the data K in the same batch is N. Select a data k, where k ∈ K ≡ {1...N}; S3-2. Call k the anchor point. Select the data of the same class as the label of k from the data K in the same batch, calculate the cosine similarity with k, and select the most similar data as p, which is used as the positive sample of k. Select the data of different classes from the label of k from the data K in the same batch, calculate the cosine similarity with k, and select the data with the lowest similarity as a, which is used as the negative sample of k; S3-3: Convert these samples into vector representations in a high-dimensional space through a graph neural network, use the Euclidean distance between vectors to measure the similarity of samples, make the distance between the anchor point and the positive example as small as possible, while making the distance between the anchor point and the negative sample as large as possible, and calculate the loss function L of contrastive learning by the following formula (h) Contrastiv : Among them, N is the number of graphs in the dataset; m k is an anchor point; m a and m k have different labels, then they are used as negative samples, and the negative sample set is represented as: A(k)≡{a∈K:m a ≠m k}; m p and m k have the same label, then they are used as positive samples, and the positive sample set is represented as: P(k)≡{p∈K:m p ≠m k}; label is a label indicating whether two samples belong to the same category. If they belong to the same category, that is, when the input is a positive sample pair, label is 0; if they belong to different categories, that is, when the input is a negative sample pair, then it is 1; m represents the minimum distance between positive and negative samples. If the distance between two samples of the same class is less than m, they are considered "similar", otherwise they are considered "dissimilar"; d is the Euclidean distance calculated between samples, that is, the distance between two samples in the feature space; S3-4. The calculation of the loss function is divided into two parts, corresponding to the same-class samples and different-class samples respectively. For the same-class samples with label 0, that is, label 0, calculate the square of the Euclidean distance to make the distances of these samples as close to 0 as possible. For the different-class samples with label 1, calculate the margin loss, that is, calculate the difference between the minimum distance threshold between the negative sample pairs and their actual Euclidean distance: m - d. When the difference is positive, that is, the distance between the negative sample pairs is too close, the loss is greater than 0. When the difference is negative, that is, the negative sample pairs are far enough away from each other, the loss is set to 0. Furthermore, by reducing the loss, the positive sample pairs are aggregated and the negative sample pairs are separated. Add the losses of the same-class samples and different-class samples and take the average value to obtain the final contrastive loss.
5. The method for detecting fake news in a graph neural network based on contrastive learning according to claim 1, wherein The specific steps of S4 include: S4-1: Aggregate the top-down propagation feature H TD and the bottom-up diffusion feature H BU obtained by the graph neural network through CGAR respectively according to the following formula (i): Among them, is the feature extracted by the last graph neural network of P-GNN; is the feature extracted by the last graph neural network of D-GNN; MEAN(*) represents the aggregation function; S4-2: Connect the aggregated propagation representation M TD and the diffusion representation M BU using the following formula (j), and merge them into the total feature M extracted by the graph neural network: M = concat(M TD , M BU ) (j); S4-3. The label y of the event passes through a fully connected layer and a Softmax layer, and the probability vectors of all classes are calculated by the following formula (k): y = Softmax(FC(M)) (k); Among them, is a probability vector for all classes used to predict event labels; S4-4: Train all parameters in the contrastive graph attention residual model for all events C by minimizing the cross-entropy between the prediction and the actual distribution Y. In the loss function, apply L2 regularization to the parameters in all graph neural network models, and calculate the classification loss function L by the following equation (m). Contrastiv : Among them, λ is the regularization coefficient; p is the CGAR model parameter; S4-5: According to the loss function L of contrastive learning Contrastive and the loss function L of classification Classification Calculate the total loss function L by the following formula (n) total : L total = L Contrastive + L Classification (n).