Microblog rumor integrated detection method based on propagation graph theme decoupling

By using the method of dissemination map theme decoupling and self-attention map pooling channel in Weibo rumors detection, the problem of low accuracy of Weibo rumors detection in the prior art is solved, and higher detection accuracy and performance are achieved.

CN120011576APending Publication Date: 2025-05-16HEBEI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510083773.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing Weibo rumor detection methods have insufficient accuracy, and they have failed to make full use of structured information during Weibo dissemination, resulting in a reduction in detection accuracy.

Method used

The Weibo rumor integrated detection method based on the theme decoupling of the propagation graph is adopted. The theme distribution of Weibo is calculated through VAETM and the theme decoupling of the Weibo propagation graph is decoupled. The graph representation learning is performed in combination with the self-attention graph pooling channel, and finally the local and global levels are integrated detection.

Benefits of technology

By learning the differences in Weibo communication process in different topics in fine-grained learning, the accuracy and performance of Weibo rumor detection are improved. The experimental results show that they are better than mainstream benchmark methods in multiple evaluation indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011576A_ABST
    Figure CN120011576A_ABST
Patent Text Reader

Abstract

The invention relates to a microblog rumor integrated detection method based on propagation graph theme decoupling. The microblog rumor integrated detection method comprises the following steps: S1, obtaining a tested microblog and text content and user information of a forwarding microblog of the tested microblog; learning text features according to the text content, and learning user features according to the user information; s2, generating a microblog propagation graph according to the forwarding relation between the tested microblog and the forwarding microblog and the text features and the user features of the microblog; s3, calculating theme distribution of the tested microblog by using VAETM, and performing theme decoupling on the microblog propagation graph; s4, performing graph representation learning on the decoupled microblog propagation graph and the decoupled microblog propagation graph by utilizing a self-attention graph pooling channel to obtain representation and global representation of the microblog propagation graph on a theme level; and S5, predicting the global representation of the microblog propagation graph and the representation of the theme level, and integrating prediction results to obtain a final prediction result. According to the method, theme decoupling is carried out on the microblogs, microblog representation is enhanced, and the accuracy of microblog prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a rumor detection method, in particular to an integrated microblog rumor detection method based on propagation graph topic decoupling. Background Art

[0002] With the rapid development of Internet technology, the speed and breadth of information dissemination have reached unprecedented levels. People can quickly obtain and disseminate information through various online platforms, but this also provides a breeding ground for the spread of rumors. The anonymity, openness and interactivity of the Internet make it easier to create and spread rumors. Anyone can post unverified information on the Internet, and this information can quickly spread to a large number of audiences, causing widespread impact. Rumor detection can identify and block the spread of rumors in the first place to avoid more serious consequences.

[0003] The graph convolutional network model for data fusion is an important technology in the existing microblog rumor detection. Its main process is divided into four modules, namely feature extraction module, feature fusion module, graph convolution module and pooling module. Specifically, the feature extraction module mainly extracts the basic information features of users, user similarity features and text features of microblogs. The feature fusion module fuses different features. After that, an undirected graph is constructed based on the forwarding between microblogs. The combined features are the feature vectors of the nodes in the graph. The graph convolution module updates the node representation through an aggregation mechanism. The pooling module combines the representations of each node to obtain the representation vector of the entire graph. Then the prediction results are obtained through a fully connected network to achieve the rumor detection task. However, this method ignores the characteristics of microblogs themselves, thereby limiting the role of the propagation process in the microblog rumor detection process and reducing the accuracy of rumor detection. Summary of the invention

[0004] The purpose of the present invention is to provide an integrated microblog rumor detection method based on propagation graph topic decoupling to solve the problem of low accuracy of rumor detection in the prior art.

[0005] The object of the present invention is achieved in that:

[0006] An integrated microblog rumor detection method based on topic decoupling of a propagation graph includes the following steps:

[0007] S1. Obtain the text content and user information of the tested microblog, as well as the text content and user information of the forwarded microblog of the tested microblog; learn text features based on the text content, and learn user features based on the user information;

[0008] S2. Generate a microblog propagation graph based on the forwarding relationship between the tested microblog and the forwarded microblog, as well as the text features and user features of the microblog;

[0009] S3. Use VAETM to calculate the topic distribution of the tested microblogs, and perform topic decoupling on the microblog propagation graph based on the topic distribution;

[0010] S4. Use the self-attention graph pooling channel to learn the graph representation of the decoupled microblog propagation graph, and obtain the representation of the decoupled microblog propagation graph at the topic level, and use the self-attention graph pooling to learn the graph representation of the microblog propagation graph before decoupling, and obtain the global representation;

[0011] S5. Predict the global representation and topic-level representation of the microblog propagation graph respectively, and integrate the prediction results to obtain the final prediction result.

[0012] Furthermore, the user information includes background information, social information and influence information;

[0013] Step S1 includes the following sub-steps:

[0014] S1-1. Based on the text content of the tested microblog and the forwarded microblog of the tested microblog, the text features are learned using the BERT model;

[0015] S1-2. Vectorize the background information and normalize the social information and the vectorized background information;

[0016] S1-3. Calculate the user's influence based on the forwarding relationship between microblogs. The calculation formula is:

[0017]

[0018] Among them, C i Represents the set of tested microblogs and their forwarded microblogs, Indicates Weibo i,j The set of forwarding edges in the φth step, e jk Indicates Weibo i,j With Weibo i,k There is a forwarding path between them;

[0019] S1-4. Use influence, normalized background information and social information as user features.

[0020] Furthermore, the specific method of generating the microblog propagation graph in step S2 is:

[0021] S2-1. According to the forwarding relationship between microblogs, the adjacency matrix A of the microblog propagation graph is calculated. The calculation formula is:

[0022]

[0023] Among them, a st is the connection relationship between node s and node t in the adjacency matrix A, εi is the forwarding relationship of Weibo, e st Indicates Weibo i,s With Weibo i,t There is forwarding between them;

[0024] S2-2. Concatenate the text features and user features of each microblog, use the concatenated features of the forwarded microblog as the parent node of the propagation graph, use the concatenated features of the forwarded microblog as the child node of the forwarded microblog, and use the concatenated features of the tested microblog as the root node of the propagation graph.

[0025] Furthermore, the specific method of calculating the topic distribution of the tested microblog in step S3 is:

[0026] S3a-1. Obtain the original microblog, input the text content of the tested microblog and the original microblog as samples into VAETM, and obtain the corresponding topic distribution;

[0027] S3a-2. Reconstruct the text of the microblog based on the obtained topic distribution;

[0028] S3a-3. Calculate the loss function of VAETM based on the reconstructed text and the initial text of the microblog, as well as the KL divergence between the generated topic distribution and the prior distribution;

[0029] S3a-4. Adjust the parameters of VAETM according to the loss function of VAETM until the loss function meets the preset conditions and output the topic distribution of the tested microblog.

[0030] Furthermore, the specific method of calculating the topic distribution corresponding to the microblog in step S3a-1 is:

[0031] S3a-1a-1. Preprocess the text content of the microblog and extract it into a vector representation. The extraction formula is:

[0032]

[0033] Among them, π i is the mixed representation of the microblog text input to VAETM, f e is a multi-layer neural network, x i is the number of times each word appears in the Weibo, is the weighted average of word vectors, is the weighted average of the word entity vectors, It is the concatenation operation of vectors;

[0034] S3a-1a-2. Use the autoencoder to obtain the mean vector μ of the microblog text i and the diagonal variance matrix The calculation formula is:

[0035] μ i =W μ π i +b μ

[0036]

[0037] Among them, W σ , b σ , W μ and b μ These are all learnable parameters in the neural network;

[0038] S3a-1a-3. Calculating the latent representation of microblogs The calculation formula is:

[0039]

[0040] Among them, ∈ (s) is an independent noise, which is an auxiliary parameter with independent marginal distribution;

[0041] S3a-1a-4. Pass the softmax function to represent the potential topic of the microblog and obtain the topic distribution θ of the tested microblog i , the calculation formula is:

[0042] S3a-1a-5. The obtained topic distribution is generated by the decoder network to generate the Dirichlet distribution η combined with the word background item i , the calculation formula is:

[0043] η i =d+θ i B

[0044] Among them, the weight matrix B obeys the Dirichlet distribution, d is the v-dimensional background vector, representing the log value of the word frequency.

[0045] Furthermore, the specific method of decoupling the microblog propagation graph in step S3 is:

[0046] The microblog diffusion graphs are copied until the number of microblog diffusion graphs is the same as the number of topics. The decoupled microblog diffusion graph set is Among them, G f This is the Weibo propagation graph corresponding to the f-th topic.

[0047] Furthermore, each self-attention graph pooling channel is stacked by three layers of attention graph pooling structures, each of which includes graph convolution, self-attention mechanism and graph pooling output;

[0048] The specific method of learning the graph representation of the decoupled microblog propagation graph is as follows:

[0049] Input the decoupled microblog propagation graph into the corresponding SAGP i The graph convolution in each layer of the attention graph pooling structure calculates the output result of the self-attention mechanism in the previous layer. Each layer of the attention graph pooling structure updates the input Weibo propagation graph by the graph convolution layer. The self-attention mechanism retains the key nodes in the updated Weibo propagation graph by calculating the self-attention scores of the nodes. The graph pooling output pools the Weibo propagation graph output by the graph convolution layer, and the output results of the three-layer attention graph pooling structure are spliced ​​and output.

[0050] Furthermore, the graph convolution aggregates the neighbor information of the node to update the node, and the graph convolution formula is:

[0051]

[0052] in, For SAGP i The input of the l-layer graph convolution, For SAGP i The initial input, X is the node representation set of the microblog propagation graph, σ(.) is the activation function; W i (l) and For SAGP i The learnable parameters in the graph convolution operation of layer l in , is the edge weight matrix of the input l-th graph convolutional layer, M (l) Normalization of M (1) =A.

[0053] Furthermore, the self-attention mechanism is used to calculate the self-attention score of the node in the updated microblog propagation graph. The calculation formula is:

[0054]

[0055] Among them, V i (l) For SAGP i The attention score vector of all nodes in the microblog propagation graph of the l-th layer attention mechanism, σ(.) is the activation function, is the normalized edge weight matrix, To pass through SAGP i The updated features of the lth convolutional layer, Θ i With b i are learnable parameters.

[0056] Furthermore, the specific way of pooling the microblog image output by the graph convolution layer is as follows:

[0057] The graph pooling output layer of each layer outputs the microblog propagation graph of the same graph convolution layer. Perform mean pooling and maximum pooling, and concatenate the results as the output of the current layer. The calculation formula is:

[0058]

[0059] Among them, s i,l for In SAGP i The output result of the lth layer in, N is The total number of nodes, x j,l for The vector representation of the jth node in is, The concatenation operation for vectors.

[0060] Early research on Weibo rumor detection mainly achieved the detection task by comprehensively analyzing the text semantics and user information of Weibo. However, Weibo text and user information only belong to Weibo itself, and it is impossible to consider the structured information contained in the propagation process into Weibo representation. To this end, recent Weibo rumor detection research has enhanced Weibo representation by learning the Weibo propagation process to further improve the performance of Weibo rumor detection. However, existing research ignores the differences in Weibo propagation process in different topics and fails to conduct fine-grained learning of Weibo propagation process from the topic level, resulting in insufficient learning of Weibo propagation process. For example, for a Weibo rumor involving two topics, health and technology, "using smartphones will lead to an increase in the incidence of Parkinson's disease", its propagation process not only involves Weibo user groups that focus on health, but also touches Weibo user groups that focus on technology. Different Weibo user groups have different acceptance, propagation speed and depth in information propagation, so the same Weibo rumor shows significant propagation differences under different topics. It can be seen that learning the Weibo propagation process in different topics can effectively avoid redundant interference between different topics, which inspires new research perspectives for using Weibo propagation process to achieve Weibo rumor detection.

[0061] The present invention performs topic decoupling on the microblog propagation graph of the tested microblog to obtain the decoupled microblog propagation graph. Through topic decoupling, the microblog propagation graph is decoupled and learned at the topic level, so as to better utilize the microblog propagation process to enhance the microblog representation. Multiple self-attention graph pooling channels are established, each self-attention graph pooling channel is trained for a topic, and the representation learning effect of the microblog propagation graph of the tested microblog is enhanced. The present invention uses a hierarchical integration mechanism to perform integrated detection of microblog rumors from both local and global levels, thereby improving the accuracy of detection. The present invention proposes an integrated detection method for microblog rumors based on topic decoupling of propagation graphs, which can perform fine-grained learning on the microblog propagation process at the topic level to further improve the performance of microblog rumor detection. Experiments are conducted on a public microblog dataset, and the results show that the proposed method is superior to the mainstream benchmark method in multiple evaluation indicators, has better microblog rumor detection performance, and provides a valuable reference for improving the performance of microblog rumor detection by optimizing microblog representation learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a flow chart of the present invention.

[0063] Figure 2 3 is a comparison chart of the detection performance of the present invention with -w / o CC and -w / o CG in rumor detection, wherein a is a comparison chart of the accuracy of the present invention with -w / o CC and -w / o CG in detecting rumors on Weibo, b is a comparison chart of the recall rate of the present invention with -w / o CC and -w / o CG in detecting rumors on Weibo, c is a comparison chart of the F1 value of the present invention with -w / o CC and -w / o CG in detecting rumors on Weibo, d is a comparison chart of the accuracy of the present invention with -w / o CC and -w / o CG in detecting non-rumors on Weibo, e is a comparison chart of the recall rate of the present invention with -w / o CC and -w / o CG in detecting non-rumors on Weibo, and f is a comparison chart of the F1 value of the present invention with -w / o CC and -w / o CG in detecting non-rumors on Weibo.

[0064] Figure 3 3 is a comparison chart of the detection performance of the present invention with -min and -mean in rumor detection, wherein a is a comparison chart of the accuracy of the present invention with -min and -mean in microblog detection of rumors, b is a comparison chart of the recall rate of the present invention with -min and -mean in microblog detection of rumors, c is a comparison chart of the F1 value of the present invention with -min and -mean in microblog detection of rumors, d is a comparison chart of the accuracy of the present invention with -min and -mean in microblog detection of non-rumors, e is a comparison chart of the recall rate of the present invention with -min and -mean in microblog detection of non-rumors, and f is a comparison chart of the F1 value of the present invention with -min and -mean in microblog detection of non-rumors.

[0065] Figure 4 3 is a comparison chart of the detection performance of rumor detection using different weights in integrated prediction, wherein a is a comparison chart of the accuracy of rumor microblog detection using different weights in integrated prediction, b is a comparison chart of the recall of rumor microblog detection using different weights in integrated prediction, c is a comparison chart of the F1 values ​​of rumor microblog detection using different weights in integrated prediction, d is a comparison chart of the accuracy of non-rumor microblog detection using different weights in integrated prediction, e is a comparison chart of the recall of non-rumor microblog detection using different weights in integrated prediction, and f is a comparison chart of the F1 values ​​of non-rumor microblog detection using different weights in integrated prediction.

[0066] Figure 5 It is a comparison chart of the accuracy of the present invention and other rumor detection methods at different stages of time and forwarding volume, wherein a is a comparison chart of the accuracy of the present invention and other rumor detection methods at different stages of time, and b is a comparison chart of the accuracy of the present invention and other rumor detection methods at different forwarding volumes. DETAILED DESCRIPTION

[0067] The present invention will be further described below in conjunction with the accompanying drawings.

[0068] like Figure 1 As shown, the integrated microblog rumor detection method based on the decoupling of the propagation graph theme of the present invention includes the following steps:

[0069] S1. Obtain the text content and user information of the tested microblog, as well as the text content and user information of the forwarded microblogs of the tested microblog; learn text features based on the text content, and learn user features based on the user information.

[0070] The present invention initializes the nodes of the microblog propagation graph according to the pre-trained BERT (Bidirectional Encoder Representations from Transformers) model, and uses the text content set T = {t0, t1, t2, ..., t n-1}, learn text features. The BERT model is a pre-trained language model based on deep learning, which can fully learn text content features and capture the contextual semantics of the text. The text features of the tested microblog and its forwarded microblog are learned through the pre-trained BERT model. The text feature set is expressed as

[0071] As shown in Table 1, since user information is a typical information for identifying rumors, the present invention selects three aspects of user information (background information, social information and influence information) to learn user characteristics. Among them, background information includes gender, Weibo age, V user certification, certification degree and location; social information includes the number of followers, the number of Weibo posts, the number of mutual fans, the number of favorites, the number of Weibo comments and the number of Weibo reposts; influence information is the user's range of influence.

[0072] Table 1: User information description

[0073]

[0074] Vectorize the statistical data of the user's background information, normalize the social information and the vectorized background information, and calculate the user's influence based on the forwarding relationship. The influence calculation formula is:

[0075]

[0076] Among them, C i represents the set of tested microblogs and their forwarded microblogs, ε φ (b i ) means microblog i The set of forwarding edges in the φth step, e ij Indicates Weibo i With Weibo j There is a forwarding path between them.

[0077] Finally got a certain Weibo b i User characteristics User characteristics specifically include influence, normalized background information, and social information as user characteristics; the user characteristics set of the tested microblog and its forwarded microblog is expressed as

[0078] S2. Generate a microblog propagation graph based on the forwarding relationship between the tested microblog and the forwarded microblog, as well as the text features and user features of the microblog.

[0079] The text features and user features corresponding to each microblog are spliced ​​together as nodes of the microblog propagation graph. The spliced ​​features of the tested microblog are used as the root node of the microblog propagation graph, and the spliced ​​features of other forwarded microblogs are child nodes; the spliced ​​features of the forwarded microblog are the parent node, and the spliced ​​features of the microblog forwarding the microblog are the child nodes. According to the forwarding relationship ε of the microblog i Define the edge relationship of the microblog propagation graph and obtain the adjacency matrix A of the microblog propagation graph. The calculation formula is:

[0080]

[0081] Among them, a st is the connection relationship between node s and node t in the adjacency matrix A, e st Indicates Weibo i,s With Weibo i,t There is forwarding between them.

[0082] The learned text features are concatenated with the user features, namely: And initialize the nodes in the microblog propagation graph. Finally, the node representation set X={X0,X1,…,X n-1 The constructed microblog diffusion graph is represented as G = <X, A>, where X is the node set of the microblog diffusion graph and A is the adjacency matrix of the microblog diffusion graph.

[0083] S3. Use VAETM to calculate the topic distribution of the tested microblogs, and perform topic decoupling on the microblog propagation graph based on the topic distribution.

[0084] Since existing research on Weibo rumor detection ignores the differences in Weibo propagation processes in different topics, the present invention performs topic decoupling on Weibo propagation graphs, which enables fine-grained learning of Weibo propagation graphs in different topics. Since the text content of most Weibo posts has a relatively concise text form and belongs to short text content, there is a lack of sufficient information to directly learn the topic distribution of Weibo posts, and it is necessary to use a topic model that can effectively learn short text content. Therefore, the present invention uses a topic model (Variational Autoencoder Topic Model, VAETM) based on variational autoencoder to learn the topic distribution of Weibo posts. VAETM enriches the semantic information of text by introducing word vectors and entity vectors, and can effectively learn the topic distribution of short texts.

[0085] VAETM is a trained model. Specifically, the text content of the original microblog and the microblog under test is input into VAETM for training. In addition, in order to prevent some noise information in the forwarded microblog from interfering with the topic learning, the model cannot accurately represent the topic distribution of the microblog under test, resulting in topic drift. The present invention learns the topic distribution of the microblog under test based on the text content of the microblog under test, and sets a threshold to determine the topic to which the microblog under test belongs. The microblog propagation graph is topic-decoupled based on the topic to which the microblog under test belongs. The process mainly includes two steps: calculating the microblog topic distribution and topic decoupling.

[0086] At least one original microblog is obtained, and the obtained original microblog and the tested microblog are input into VAETM as training samples for training. When the number of microblogs used for training reaches more than 1,000, a more appropriate number of topics and topic distribution can be obtained. First, the specific number of topics is determined based on the perplexity and text consistency.

[0087] When calculating the distribution of microblog topics, there are two main steps: the inference process and the generation process. The inference process is used to calculate the potential topic representation of microblogs, and the generation process restores the text of microblogs through the potential topic representation, and trains the model by comparing the difference between the restored text and the initial text.

[0088] The reasoning process mainly calculates the potential topic representation of Weibo through the text content of Weibo. First, the text content of all Weibo input into VAETM is preprocessed and extracted into vector representation. The extraction formula is:

[0089]

[0090] Among them, π i is the mixed representation of the microblog text input to VAETM, f e is a multi-layer neural network, x i is the number of times each word appears in the Weibo, is the weighted average of word vectors, is the weighted average of the word entity vectors, It is the concatenation operation of vectors.

[0091] Secondly, we assume that the latent topic representation of Weibo follows a normal prior distribution and use an autoencoder to obtain the mean vector μ of the Weibo text i and the diagonal variance matrix The calculation formula is:

[0092] μ i =W μ π i +b μ

[0093]

[0094] Among them, W σ , b σ , W μ and b μ are all learnable parameters in the neural network, μ i is the mean vector, is the diagonal variance matrix.

[0095] Finally, using the resampling technique, we can get the mean vector μ i , diagonal variance matrix With independent noise ∈ (s) Get the potential topic representation of Weibo The calculation process is:

[0096]

[0097] Among them, ∈ (s)is an auxiliary parameter with independent marginal distributions.

[0098] After obtaining the latent topic representation of the microblog, it is necessary to further obtain the topic distribution of the microblog through the generation process, and restore the text content of the microblog through the topic distribution. First, the latent topic representation of the microblog obtained in the inference process is The topic distribution θ of Weibo is obtained through the softmax function i , and its calculation formula is:

[0099] The obtained topic distribution is passed through the decoder network to generate a Dirichlet distribution η combined with the word background term i , the calculation formula is:

[0100] η i =d+θ i B

[0101] Among them, the weight matrix B obeys the Dirichlet distribution, d is the v-dimensional background vector, representing the log value of the word frequency.

[0102] Then, the initial text of the microblog is restored by extracting words through polynomials. The calculation formula is:

[0103] w ij ~Mult(η i )

[0104] Among them, w ij is the word set of the restored text, and Mult(.) represents multinomial distribution.

[0105] The loss function of VAETM consists of the reconstruction loss between the reconstructed text and the initial text and the KL divergence between the generated topic distribution and the prior distribution. The formula is as follows:

[0106]

[0107] in, is the reconstruction loss, N i is the total number of words in microblog i, c is the regular term coefficient, is the KL divergence, is the posterior distribution of the hypothesis, w i For Weibo i,0 The word set in is the true posterior distribution, and α represents the hyperparameter of the Dirichlet distribution.

[0108] The VAETM is trained through the loss function, and the parameters in the VAETM are changed to finally obtain the topic distribution of the tested microblogs. The expression is:

[0109]

[0110] Among them, k is the total number of topics, z i is the i-th topic, p i is the probability that a microblog belongs to the i-th topic.

[0111] In order to eliminate the interference of some noise in the process of calculating the distribution of the tested microblog topics, it is necessary to further determine the topic of the tested microblog so as to better decouple the microblog propagation graph. Z To determine the topic of the tested microblog where z j represents the jth topic. The threshold is set based on the total number of topics k. The probability of the topic distribution being greater than θ Z Add the topic of the tested microblog to the collection of topics to obtain the collection of topics to which the tested microblog belongs Decouple the microblog diffusion graph according to the topic of the tested microblog, copy the microblog diffusion graph until the number of microblog diffusion graphs is the same as the number of topics, and obtain the decoupled microblog diffusion graph set Among them, G j It is the microblog propagation graph corresponding to the jth topic.

[0112] S4. Use self-attention graph pooling to learn the graph representation of the decoupled Weibo propagation graph to obtain the representation of the decoupled Weibo propagation graph at the topic level, and use self-attention graph pooling to learn the graph representation of the Weibo propagation graph before decoupling to obtain the global representation.

[0113] The present invention establishes multiple self-attention graph pooling channels to perform fine-grained learning on the decoupled Weibo propagation graph. Self-Attention Graph Pooling (SAGP) is a graph convolutional neural network method based on an improved self-attention mechanism. It retains the key nodes in the propagation graph through the self-attention mechanism, and optimizes the pooling process of the graph convolutional neural network to enhance the representation learning effect of the propagation graph. In addition, in order to prevent the detection method from over-learning and relying on the representation of the Weibo propagation graph at the topic level, resulting in overfitting. The present invention retains the Weibo propagation graph G before decoupling to learn the global representation of the Weibo propagation graph, and further improves the accuracy of the detection method by comprehensively considering the global representation of the Weibo propagation graph and its representation at the topic level. The Weibo propagation graph G before decoupling and the Weibo propagation graph after decoupling are combined. The input is sent to the corresponding self-attention map pooling channel for fine-grained learning. The set of self-attention map pooling channels is represented as {SAGP1, SAGP2, ..., SAGP k ,SAGPa}. Among them, {SAGP1,SAGP2,...,SAGP k ,SAGP a}Representation learning for the decoupled microblog propagation graph set, SAGP a It is a global self-attention graph pooling channel, which performs representation learning on the microblog propagation graph G before decoupling. All self-attention graph pooling channels adopt a layered architecture, which is composed of three layers of attention graph pooling structures. Each layer of attention graph pooling structure consists of graph convolution, self-attention mechanism and graph pooling output.

[0114] The decoupled microblog propagation graph is used for graph representation learning. Specifically, the decoupled microblog propagation graph is input into the corresponding SAGP i The graph convolution in each layer of the attention graph pooling structure calculates the output result of the self-attention mechanism in the previous layer. Each layer of the attention graph pooling structure updates the input Weibo propagation graph by the graph convolution layer. The self-attention mechanism retains the key nodes of the updated Weibo propagation graph. The graph pooling output pools the Weibo propagation graph output by the graph convolution layer. The output results of the three-layer attention graph pooling structure are spliced ​​and output.

[0115] First, the decoupled microblog propagation graph G i Input to SAGP i , update the node by aggregating the neighbor information of the node through graph convolution to capture the potential propagation characteristics. The graph convolution formula is:

[0116]

[0117] in, For SAGP i The input of the l-layer graph convolution, For SAGP i The initial input, X is the node representation set of the microblog propagation graph, σ(.) is the activation function; W i (l) and For SAGP i The learnable parameters in the graph convolution operation of layer l in . is the edge weight matrix of the input l-th graph convolutional layer, M (l) Normalization of M (1) =A, the matrix normalization process specifically includes adding the unit matrix After normalization, for The degree matrix of .

[0118] Through the graph convolution operation, the decoupled microblog propagation graph G i After the update, the self-attention mechanism is used to calculate the self-attention score of the node and retain the key node and node relationship. The self-attention score is calculated by the graph convolution method, and the formula is:

[0119]

[0120] in, is the attention score vector of all nodes in the microblog propagation graph, is the normalized edge weight matrix, H i ` is the feature updated by the graph convolution layer, Θ i With b i are learnable parameters.

[0121] After obtaining the self-attention score of each node, a pooling rate ρ∈(0,1] is set through the top-rank method, and the self-attention score is selected. Nodes are used as key nodes in the Weibo propagation graph:

[0122]

[0123] Among them, top-rank() is used to return the top ranking with the highest score in V. Index value, V mask is the self-attention mask, idx For index operations.

[0124] After the self-attention mask is used to obtain the retained key nodes, the formula process is as follows:

[0125]

[0126] in, For the node input of the next layer, is the characteristic matrix indexed by row, ⊙ represents the dot product calculation of the elements in the two matrices one by one, M (l+1) is the edge weight matrix of the next layer, A idx,idx is an adjacency matrix indexed by rows and columns.

[0127] The graph pooling output layer of each layer outputs the microblog propagation graph H of the same graph convolution layer. i `Perform mean pooling and maximum pooling, and concatenate the results as the output of the current layer. The output formula is:

[0128]

[0129] Where N is the total number of nodes, x j,l represents the updated vector representation of the jth node at layer l, Represents a vector concatenation operation.

[0130] Finally, the output results of the three layers are combined to obtain the decoupled microblog propagation graph G i In SAGP i The representation s learned in i , the calculation process is:

[0131]

[0132] Among them, s i,l Indicated in SAGP i The output result of the lth layer in is, Represents a vector concatenation operation.

[0133] Similarly, the microblog propagation graph G is passed through the global self-attention graph pooling channel SAGP a Get the global representation s of the microblog propagation graph G a , the calculation process is:

[0134]

[0135] Among them, s a,l Indicated in SAGP a The output result of the lth layer in is, Represents a vector concatenation operation.

[0136] S5. Predict the global representation and topic-level representation of the microblog propagation graph respectively, and integrate the prediction results to obtain the final prediction result.

[0137] The present invention uses multiple independent multi-layer perceptrons to predict the microblog propagation graph representations obtained by different self-attention graph pooling channels respectively. The set of multi-layer perceptrons is represented as {MLP1, MLP2, ..., MLP k ,MLP a}, where {MLP1,MLP2,...,MLP k} For the decoupled microblog propagation graph in {SAGP1,SAGP2,...,SAGP k} to predict the set of representations obtained in the topic level of microblogs. a The microblog propagation graph before decoupling in SAGP k In addition, the present invention designs a hierarchical integration mechanism, which first integrates the prediction results of Weibo at the topic level as local prediction results, and then further integrates the local prediction results with the global prediction results, so as to realize the integrated detection of Weibo rumors from both the local and global levels, so as to further improve the accuracy of the detection method.

[0138] Post the Weibo communication map on SAGP i The representation s obtained in i Through MLP i The prediction results of Weibo in topic i are:

[0139]

[0140] Among them, W i and are learnable parameters.

[0141] According to the topic of the microblog propagation graph, the corresponding multi-layer perceptron obtains the prediction result set At the same time, these prediction results are preliminarily integrated to obtain local prediction results. Generally speaking, when a microblog is not a rumor, it should be non-rumor in all topics. Therefore, if this microblog is identified as a rumor in a certain topic, it can be directly regarded as a rumor. Therefore, the present invention takes the prediction result with the highest rumor probability in the prediction result set as the local prediction result. Local prediction results The calculation process is:

[0142]

[0143] Among them, max(.) is the maximum value function, For Weibo in MLP j The prediction results obtained in .

[0144] At the same time, the Weibo communication map will be posted on SAGP a The representation s obtained in a Through MLP a Get the global prediction result y of the Weibo propagation graph a , the calculation process is:

[0145]

[0146] Among them, W a and b a are learnable parameters.

[0147] In order to comprehensively consider the local prediction results And the global prediction results The present invention sets a fusion weight γ to further integrate the local prediction results with the global prediction results to obtain the final prediction result. The calculation process is:

[0148]

[0149] Finally, based on the final prediction results To determine whether a Weibo post is a rumor and realize the Weibo rumor detection task.

[0150] The present invention performs integrated detection on microblog rumors from two levels, local and global, to achieve the microblog rumor detection task more accurately.

[0151] The self-attention map pooling channel, the global self-attention map pooling channel and the multi-layer perceptron are all obtained through training. For the self-attention map pooling channel, the training samples are the Weibo propagation graphs of the original Weibo with known detection results and decoupled. For the global self-attention map pooling channel, the training samples are the Weibo propagation graphs of the original Weibo with known detection results, and whether it is a rumor is used as the label of the sample.

[0152] The topics of the original microblogs trained in the self-attention map pooling channel are the same as those of the tested microblogs.

[0153] The loss function used in the training process of the present invention is to minimize the cross entropy loss. k Channel and MLP1~MLP k , and its corresponding cross entropy loss set is in Defined as:

[0154]

[0155] in, For SAGP i and MLP i The cross entropy loss is For training SAGP i Weibo collection, y j is the true result of the jth original microblog, is the prediction result of the jth original microblog in topic i.

[0156] For SAGP a and MLP a , whose cross entropy loss The calculation formula is:

[0157]

[0158] in, For training SAGP a Collection of Weibo, is the global prediction result of the jth original microblog.

[0159] S5. Method evaluation.

[0160] As shown in Table 2, the present invention uses the microblog dataset published in the study of Ma et al. to evaluate the effectiveness of the method. The dataset contains 2313 rumors and 2351 non-rumors. The present invention performs data statistics on the dataset and lists the detailed statistical results of the dataset in .

[0161] Table 2: Microblog data statistics

[0162] Statistics rumor Not a rumor Original number of microblogs 2 313 2 351 Number of users 2 025 513 1 619 219 Number of retweets 2 088 430 1 659 152 Minimum forwarding quantity 3 5 Maximum forwarding number 59 317 52 156 Average number of forwarding 903 706

[0163] The present invention uses common evaluation indicators in rumor detection problems, namely, accuracy, precision, recall and F1 value, to evaluate the performance of the microblog rumor detection method. The calculation formula of the relevant indicators is as follows:

[0164]

[0165] Among them, P represents the number of rumors, N represents the number of non-rumors, TP represents the number of correctly identified rumors, TN represents the number of correctly identified non-rumors, FP represents the number of non-rumors identified as rumors, and FN represents the number of rumors identified as non-rumors. Generally speaking, the higher the accuracy, precision, recall and F1 value, the better the performance of the detection method.

[0166] The text selects a Weibo rumor detection method based on traditional artificially constructed features, a Weibo rumor detection method based on kernel functions, and a Weibo rumor detection method based on deep learning as benchmark methods, and compares them with the present invention (i.e., GTD-SAGP).

[0167] DTR: This method is based on decision tree ranking and realizes the identification of rumor events through regular expression matching of rumor signal features.

[0168] DTC: This method manually extracts features and keywords to obtain the confidence level of news and predicts rumors based on the confidence level.

[0169] SVM-TS: This method uses the SVM classifier of time series to model rumor representation by manually extracting features.

[0170] SVM-HK: This method uses a hybrid kernel function formed by a random walk graph and an RBF kernel to simulate the spread of rumors, taking advantage of the propagation characteristics of rumors.

[0171] RvNN: This method uses recurrent neural networks to model the propagation structure of rumors and optimizes the learning of rumor propagation characteristics.

[0172] RDEA: This method is based on graph convolutional networks. It enriches the semantic information of rumors through contrastive learning and event expansion, and integrates the propagation characteristics and text characteristics of rumors.

[0173] Bi-GCN: This method uses graph convolutional networks to learn complex rumor propagation features through a bidirectional propagation structure to achieve higher detection effects.

[0174] DA-GCN: This method uses a dual attention mechanism and a graph convolutional network to fuse the text features and propagation features of rumors, learn complex representations of rumors, and perform rumor detection.

[0175] Table 2: Rumor detection performance of different methods in Weibo dataset

[0176]

[0177] The bold values ​​represent the optimal values ​​of the evaluation indicators, and the underlined values ​​represent the suboptimal values.

[0178] It can be observed from Table 2 that the detection effect of the present invention is better than other methods in terms of accuracy, recall rate and F1 value. Compared with other methods, GTD-SAGP performs fine-grained learning on the microblog propagation process from the topic level, thereby enhancing the representation effect of the microblog propagation graph, and through hierarchical integration, comprehensively considers the global and local prediction results of the microblog propagation graph to improve the microblog rumor detection performance. The experimental results prove that the present invention has better microblog rumor detection performance than other methods.

[0179] This paper uses ablation experiments to test whether the design of each part in the GTD-SAGP method is effective, so as to prove the effectiveness and rationality of the method proposed in the text. The ablation experiments are mainly conducted on the two parts of the propagation graph topic decoupling and hierarchical integration mechanism.

[0180] Based on the original method, the present invention generates two variants, namely -w / o CC and -w / o CG. Compared with the original method, -w / o CC removes the process of learning the representation of the Weibo diffusion graph at the topic level and only learns the global representation of the Weibo diffusion graph. -w / o CG removes the process of learning the global representation of the Weibo diffusion graph and only learns the representation of the Weibo diffusion graph at the topic level. By comparing the detection performance among the three, the effectiveness of the topic decoupling of the diffusion graph is verified.

[0181] like Figure 2 As shown in the figure, the original method outperforms other variant methods in multiple evaluation indicators, and -w / o CG outperforms -w / o CC in overall performance. The experimental results show that the topic decoupling of the propagation graph helps to enhance the effect of learning the representation of the microblog propagation graph, thereby improving the accuracy of detection.

[0182] The present invention generates two variants according to the hierarchical integration mechanism of the proposed method, namely -min and -mean. Among them, -min takes the prediction result with the lowest rumor probability in the microblog propagation graph among different topics as the local prediction result. And -mean averages the prediction results of the microblog propagation graph in different topics through mean calculation, and takes the prediction results after averaging as the local prediction result. The hierarchical integration mechanism of the present invention takes the prediction result with the highest rumor probability in the microblog propagation graph among different topics as the local prediction result.

[0183] like Figure 3 As shown in the figure, the hierarchical integration mechanism of the method proposed in the present invention achieves the best effect, -mean is second, and -min is the worst. The experimental results verify the rationality of the hierarchical integration mechanism.

[0184] In order to comprehensively consider the global prediction results and local prediction results of the microblog propagation graph, the present invention sets a weight γ to integrate them, and determines the value of γ through experiments.

[0185] like Figure 4 As shown in the figure, when the fusion weight is set to 0.6, all evaluation indicators in rumors and non-rumors are optimal, indicating that both global prediction results and local prediction results have an important impact on detection. At the same time, it can be observed that the detection effect when the weight is zero is much lower than the detection effect when the weight is one. This shows that local prediction results are more helpful in achieving rumor detection tasks than global prediction results.

[0186] Early rumor detection is to discover and identify rumors in the early stages of rumor propagation, intervene in time, and reduce the harm caused by rumors. At the same time, this is also an important indicator for evaluating detection performance. The present invention evaluates the early rumor detection capabilities of the method proposed in the present invention and other baseline methods by limiting the number of reposts of the tested microblog and the time elapsed after its release. The present invention selects a microblog rumor detection method based on deep learning that ranks high in performance in comparative experiments for comparison.

[0187] like Figure 5 As shown in the figure, at different stages of time, the method proposed by the present invention can achieve an accuracy rate of about 93% 1 hour after the microblog is released, and the detection accuracy rate in the entire microblog propagation process is better than other benchmark methods. In the detection of microblog rumors with different forwarding amounts, the present invention can achieve an accuracy rate of about 93% after the number of forwarding exceeds 10. The experimental results show that compared with other benchmark methods, the present invention can not only achieve a higher accuracy rate, but also realize the early detection of microblog rumors in a shorter time.

[0188] Dataset: Ma J, Gao W, Mitra P, et al. Detecting rumors from microblogs with recurrent neural networks[C] / / In Proceedings of the 25th International Joint Conference on Artificial Intelligence (IJCAI), New York, 2016: 3818-3824.

[0189] Comparison methods:

[0190] DTR: Mihalcea R, Strapparava C. The lie detector: Explorations in the automatic recognition of deceptive language[C] / / In Proceedings of the ACL-IJCNLP 2009 Conference Short Papers, Singapore, 2009: 309-312.

[0191] DTC: Castillo C, Mendoza M, Poblete B. Information credibility on Twitter[C] / / In Proceedings of the 20th Int Conf on World Wide Web, Berlin, 2011: 675-684.

[0192] SVM-TS: Wu K, Yang S, Zhu K Q. False rumors detection on sina weibo by propagation structures[C] / / 2015 IEEE 31st International Conference on Data Engineering, Korea (South), 2015: 651-662.

[0193] SVM-HK:Ma J,Gao W,Wei Z Y,et al.Detect rumors using time series ofsocial context information on microblogging websites[C] / / In Proceedings ofthe 24th ACM International on Conference on Information and KnowledgeManagement,New York,2015:1751-1754.

[0194] RvNN:Ma J,Gao W,Wong K F.Rumor detection on twitter with tree-structured recursive neural networks[C] / / In Proceedings of the 56th AnnualMeeting of the Association for Computational Linguistics,Melbourne,2018:1980-1989.

[0195] REDA:Yang Y J,Wang Y H,Wang L,et al.PostCom2DR:Utilizing informationfrom post and comments to detect rumors[J].Expert Systems with Applications,2022,189:116071.

[0196] Bi-GCN:Bian T,Xiao X,Xu T Y,et al.Rumor detection on social mediawith bi-directional graph convolutional networks[C] / / In Proceedings of theAAAI Conference on Artificial Intelligence,2020:549-556.

[0197] DA-GCN:Liu X Y,Zhao Z Y,Zhang Y H,et al.Social Network RumorDetection Method Combining Dual-Attention Mechanism With Graph ConvolutionalNetwork[J].IEEE Transactions on Computational Social Systems,2023,10(5):2350-2361.

Claims

1. An integrated microblog rumor detection method based on the decoupling of the propagation graph, characterized in that: The steps include: S1. Obtain the text content and user information of the tested microblog, as well as the text content and user information of the forwarded microblog of the tested microblog; learn text features based on the text content, and learn user features based on the user information; S2. Generate a microblog propagation graph based on the forwarding relationship between the tested microblog and the forwarded microblog, as well as the text features and user features of the microblog; S3. Use VAETM to calculate the topic distribution of the tested microblogs, and perform topic decoupling on the microblog propagation graph based on the topic distribution; S4. Use the self-attention graph pooling channel to learn the graph representation of the decoupled microblog propagation graph, and obtain the representation of the decoupled microblog propagation graph at the topic level, and use the self-attention graph pooling to learn the graph representation of the microblog propagation graph before decoupling, and obtain the global representation; S5. Predict the global representation and topic-level representation of the microblog propagation graph respectively, and integrate the prediction results to obtain the final prediction result.

2. The integrated microblog rumor detection method according to claim 1, characterized in that: The user information includes background information, social information and influence information; Step S1 includes the following sub-steps: S1-1. Based on the text content of the tested microblog and the forwarded microblog of the tested microblog, the text features are learned using the BERT model; S1-2. Vectorize the background information and normalize the social information and the vectorized background information; S1-3. Calculate the user's influence based on the forwarding relationship between microblogs. The calculation formula is: Among them, C i Represents the set of tested microblogs and their forwarded microblogs, Indicates Weibo i,j The set of forwarding edges in the φth step, e jk Indicates Weibo i,j With Weibo i,k There is a forwarding path between them; S1-4. Use influence, normalized background information and social information as user features.

3. The integrated microblog rumor detection method according to claim 1 is characterized in that: The specific method of generating the microblog propagation graph in step S2 is: S2-1. According to the forwarding relationship between microblogs, the adjacency matrix A of the microblog propagation graph is calculated. The calculation formula is: Among them, a st is the connection relationship between node s and node t in the adjacency matrix A, ε i is the forwarding relationship of Weibo, e st Indicates Weibo i,s With Weibo i,t There is forwarding between them; S2-2. Concatenate the text features and user features of each microblog, use the concatenated features of the forwarded microblog as the parent node of the propagation graph, use the concatenated features of the forwarded microblog as the child node of the forwarded microblog, and use the concatenated features of the tested microblog as the root node of the propagation graph.

4. The integrated microblog rumor detection method according to claim 1, characterized in that: The specific method of calculating the topic distribution of the tested microblog in step S3 is: S3a-1. Obtain the original microblog, input the text content of the tested microblog and the original microblog as samples into VAETM, and obtain the corresponding topic distribution; S3a-2. Reconstruct the text of the microblog based on the obtained topic distribution; S3a-3. Calculate the loss function of VAETM based on the reconstructed text and the initial text of the microblog, as well as the KL divergence between the generated topic distribution and the prior distribution; S3a-4. Adjust the parameters of VAETM according to the loss function of VAETM until the loss function meets the preset conditions and output the topic distribution of the tested microblog.

5. The integrated microblog rumor detection method according to claim 4 is characterized in that: The specific method of calculating the topic distribution corresponding to the microblog in step S3a-1 is: S3a-1a-1. Preprocess the text content of the microblog and extract it into a vector representation. The extraction formula is: Among them, π i is the mixed representation of the microblog text input to VAETM, f e is a multi-layer neural network, x i is the number of times each word appears in the Weibo, is the weighted average of word vectors, is the weighted average of the word entity vectors, Represents a vector concatenation operation; S3a-1a-2. Use the autoencoder to obtain the mean vector μ of the microblog text i and the diagonal variance matrix The calculation formula is: m i =W μ p i +b μ Among them, W σ 、b σ , W μ and b μ These are all learnable parameters in the neural network; S3a-1a-3. Computing latent topic representations of microblogs The calculation formula is: Among them, ∈ (s) is an independent noise, which is an auxiliary parameter with independent marginal distribution; S3a-1a-4. Pass the softmax function to represent the potential topic of the microblog and obtain the topic distribution θ of the tested microblog i , the calculation formula is: S3a-1a-5. The obtained topic distribution is generated by the decoder network to generate the Dirichlet distribution η combined with the word background item i , the calculation formula is: or i =d+θ i B Among them, the weight matrix B obeys the Dirichlet distribution, d is the v-dimensional background vector, representing the log value of the word frequency.

6. The integrated microblog rumor detection method according to claim 1, characterized in that: The specific method of decoupling the microblog propagation graph in step S3 is: The microblog diffusion graphs are copied until the number of microblog diffusion graphs is the same as the number of topics. The decoupled microblog diffusion graph set is Among them, G f This is the Weibo propagation graph corresponding to the f-th topic.

7. The integrated microblog rumor detection method according to claim 1, characterized in that: Each self-attention graph pooling channel is composed of three layers of attention graph pooling structures, each of which includes graph convolution, self-attention mechanism and graph pooling output; The specific method of learning the graph representation of the decoupled microblog propagation graph is as follows: Input the decoupled microblog propagation graph into the corresponding SAGP i The graph convolution in each layer of the attention graph pooling structure calculates the output result of the self-attention mechanism in the previous layer. Each layer of the attention graph pooling structure updates the input Weibo propagation graph by the graph convolution layer. The self-attention mechanism retains the key nodes in the updated Weibo propagation graph by calculating the self-attention scores of the nodes. The graph pooling output pools the Weibo propagation graph output by the graph convolution layer, and the output results of the three-layer attention graph pooling structure are spliced ​​and output.

8. The integrated microblog rumor detection method based on the decoupling of the propagation graph according to claim 7 is characterized in that: The graph convolution aggregates the neighbor information of the node and updates the node. The graph convolution formula is: in, For SAGP i The input of the l-layer graph convolution, For SAGP i The initial input, X is the node representation set of the microblog propagation graph, σ(.) is the activation function; W i (l) and For SAGP i The learnable parameters in the graph convolution operation of layer l in , is the edge weight matrix of the input l-th graph convolutional layer, M (l) Normalization of M (1) =A.

9. The integrated microblog rumor detection method based on the decoupling of the propagation graph according to claim 7 is characterized in that: The self-attention mechanism is used to calculate the self-attention score of the node in the updated Weibo propagation graph. The calculation formula is: Among them, V i (l) For SAGP i The attention score vector of all nodes in the microblog propagation graph of the l-th layer attention mechanism, σ(.) is the activation function, is the normalized edge weight matrix, To pass through SAGP i The updated features of the lth convolutional layer, Θ i With b i are learnable parameters.

10. The integrated microblog rumor detection method based on the decoupling of the propagation graph according to claim 7 is characterized in that: The specific way of pooling the microblog image output by the graph convolution layer is as follows: The graph pooling output layer of each layer outputs the microblog propagation graph of the same graph convolution layer. Perform mean pooling and maximum pooling, and concatenate the results as the output of the current layer. The calculation formula is: Among them, s i,l for In SAGP i The output result of the lth layer in, N is The total number of nodes, x j,l for The vector representation of the jth node in is, This is the concatenation operation for vectors.