Rumor detection method based on information propagation structure

By using the deep wandering spanning tree algorithm and user forwarding network in rumors detection, combined with the Graph Transformer model, the problems of insufficient early detection and dissemination feature extraction and limitations of dissemination structure modeling are solved, and more accurate and richer rumor detection effects are achieved.

CN120216780AActive Publication Date: 2025-06-27SOUTHEAST UNIV +1

Patent Information

Application Number
CN202510270076.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-04
Filing Date
2025-03-07
Publication Date
2025-06-27
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The prior art is difficult to achieve early detection, insufficient extraction of propagation feature and limitations of propagation structure modeling in rumor detection, resulting in poor detection results.

Method used

Using a rumor detection method based on information dissemination structure, the deep wandering spanning tree algorithm and user forwarding network are designed, structural encoding and location encoding are extracted, and multi-dimensional information characteristics are integrated for rumor detection.

Benefits of technology

It improves the accuracy and richness of rumor detection, can better capture the complexity of information dissemination and user behavior characteristics, and achieve more accurate rumor identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216780A_ABST
    Figure CN120216780A_ABST
Patent Text Reader

Abstract

The rumor detection method based on the information propagation structure comprises the following steps: constructing a tweet propagation tree by social network data according to a forwarding relationship; converting the propagation tree into a user forwarding sub-graph through a node replacement strategy, and constructing a global user forwarding network through node merging and edge attribute fusion; respectively extracting local and global structure features of the user forwarding and pushing network by using a deep walk spanning tree algorithm and a WL algorithm, and splicing the local and global structure features to obtain a complete structure code; meanwhile, tweet content features are extracted through a pre-trained large language model to obtain content codes, position code information on a user forwarding and pushing network is calculated through algorithms such as a Laplacian matrix and intimacy sorting, and then the corresponding codes are spliced and input into a Graph Transform model to capture association between the features and the authenticity of propagation tweet. And finally, judging whether the tweet is a rumor or not by using an MLP classifier. And an efficient and reliable solution is provided for rumor detection in the social network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of rumor detection, and specifically to a rumor detection method based on the information dissemination structure. Background Art

[0002] With the rapid development of Internet technology and mobile Internet technology, online social media platforms have gradually become important channels for information dissemination and acquisition in the new media environment. In recent years, the number of users and content on social platforms such as Twitter and Sina Weibo has increased sharply, and online social network services have become increasingly perfect. Compared with traditional media dissemination methods such as television and newspapers, online social media platforms have certain advantages in terms of real-time nature, freedom, dissemination speed, user interaction, etc., and have received extensive attention and use from users.

[0003] However, due to the strong real-time nature of the information published on social media platforms but relatively lagging platform reviews, some unverified or deliberately fabricated information can often spread rapidly in a short period of time, which provides a breeding ground for the emergence and large-scale spread of rumors. These rumors usually not only limit to misleading readers to generate wrong cognitions or negative emotions, but also further incite social confrontation and exacerbate the deterioration of the situation. Therefore, we hope to identify possible rumors through timely and effective detection means, so that the platform's supervisors can guide and control them.

[0004] The first challenge is the early detection of rumors. At present, after a speech that may involve a rumor is published, it is very difficult to intervene in a timely and effective non-artificial manner before the influence of this information spreads widely. Only some emergencies or hot social topics will be subject to additional reviews. If potential rumor information can be detected early and it can be predicted that this information may have a huge impact on the public in a period of time in the future based on its high-frequency click volume in the short term, then the platform can label it in time, remind the majority of Internet users to judge carefully, report it upwards, and intervene and manage it in time to avoid adverse social impacts. Therefore, it is very necessary to effectively detect rumors early.

[0005] The second challenge is the insufficient extraction of propagation features. The spread of rumors often depends on the social relationships and interaction patterns among users, such as forwarding and commenting. Some users may act as "spreaders" to quickly spread rumors to a wider range, while others may inhibit the spread of rumors through comments or forwards. However, if these complex interaction relationships cannot be effectively modeled, the detection model may ignore key social network features, thereby reducing the detection effect. These methods usually regard each tweet and its propagation tree as an independent graph structure, ignoring the associations between different propagation trees. In addition, existing interaction feature extraction methods are difficult to effectively model the relationships between tweets posted and forwarded by the same user in multiple propagation trees. This independent processing method causes the model to fail to capture the information associations across propagation trees, thereby reducing the accuracy of rumor detection. The insufficient extraction of interaction features directly affects the effect of rumor detection. First of all, the behavior of users posting and forwarding tweets in multiple propagation trees often contains rich information, which is of great significance for judging their motives and behavior patterns. If these relationships cannot be effectively modeled, the model will not be able to make full use of the user's historical behavior data, thereby reducing the robustness of the detection algorithm. Secondly, the associations between different propagation trees can reveal the potential laws and network structure features of rumor propagation, which is of great value for identifying complex rumor propagation mechanisms. In addition, the lack of interaction features may also cause the model to be unable to distinguish between real information and rumors, especially when faced with rumors with high ambiguity or concealment. This shows the importance of deeply mining rumor propagation features.

[0006] The third challenge is the limitation of propagation structure modeling. Current graph structure representation learning techniques face certain limitations in the rumor detection task, especially in capturing the rumor propagation structure features. Early dimension reduction-based methods can handle nonlinear data, but lack scalability and are prone to overfitting. While matrix factorization-based methods provide a global perspective, they have insufficient ability to distinguish local structure differences and are difficult to meet the requirements of rumor detection for local feature sensitivity. In addition, classic random walk-based methods are inefficient in processing graph structures composed of propagation trees, have limited information discrimination for nodes and edges, and are difficult to deeply explore the deep structure of the graph. In contrast, deep learning-based methods can handle complex graph data, but still have obvious shortcomings in capturing long-distance dependencies of nodes, using edge attribute information, and processing large-scale directed graph data. Early methods mainly described these interaction relationships by constructing propagation trees, but this method is difficult to handle complex network structures and multi-type data. In addition, traditional RvNN (Random Walk Neural Network) or GCN (Graph Convolutional Network) can capture some structure features, but show high computational complexity and limited adaptability when faced with large-scale and dynamically changing social networks. To solve this problem, it is necessary to effectively model and extract the discriminative features between rumors and real information in the dynamic propagation process.

[0007] For the rumor detection task on online social networks, following the development context of research methods, it can be mainly divided into: detection methods mainly based on traditional machine learning methods and detection methods mainly based on deep learning methods. The research on traditional machine learning methods mainly focuses on the exploration of static attributes of tweets themselves and in the propagation process in feature engineering. In the research on information credibility on platform X, Castillo et al. systematically proposed four directions of features, namely, text content-based methods, user attribute-based methods, topic-based methods, and propagation-based methods, making a pioneering contribution to the research on information credibility. In the same year, Qazvinian proposed features related to URLs and hashtags associated with tweets for platform features, such as calculating the distribution of topic hashtag labels in normal information and rumors respectively. Wu modeled the propagation process of information as a tree structure and used the graph kernel algorithm based on random walk to model the similarity between propagation trees. Based on this, combined with message features, user features, and forwarding features, methods such as SVM were used to achieve rumor detection.

[0008] The deep learning-based methods mainly focus on modeling the context of the information propagation process and mining the dynamic semantic representation of text. Graves et al. proposed the LSTM model to alleviate the long-term dependency problem that is difficult to solve in RNN networks. In 2014, Chung et al. simplified the gated units in LSTM and proposed the GRU model, which greatly reduced the number of model parameters and improved the model's calculation rate. At present, the LSTM model and the GRU model are the most widely used RNN sequence models. In 2016, Ma et al. simplified the information propagation process into a simple linear process according to the time relationship at the international top conference IJCAI2016 for the first time, thereby simulating the context of information propagation. For the first time, the LSTM deep model was used to model the information sequence, capture the information dependency relationship at adjacent moments, and perform rumor detection. On this basis, Ling et al. proposed a multi-task rumor detection model. On the one hand, the LSTM model is used to detect the position and attitude of each user towards the original tweet. On the other hand, the rumor detection is realized by combining the user's text content and user information. The rumor detection method based on the RNN recurrent neural network ignores the interactive relationship of information in the real propagation process. Ling et al. improved the mask mechanism of Transformer so that it can process graph structure data that simulates the information propagation process. The parallel mechanism of Transformer greatly improves the efficiency of calculation. Based on the research of Ma et al., Bian et al. used Bi-GCN technology to simultaneously process the information propagation and diffusion process in two directions, and finally integrated the information in different directions to achieve rumor detection. On this basis, Yang et al. considered the above relationships at the same time, constructed a heterogeneous network with users and tweets as nodes, and realized rumor detection on this network through adversarial learning. In addition, some studies have also introduced large language models (LLM) for application in tasks.

[0009] In summary, detection methods based on traditional machine learning usually rely on the researchers' prior knowledge and it is difficult to ensure the completeness of features. From the perspective of content features, simple statistical features are difficult to represent potential high-level semantic features and abstract information such as opinions and emotions. From a higher level, it is difficult for different types of features to interact with each other, and it is impossible to effectively integrate different types of features. Detection methods based on deep learning, whether they are RNN series methods, GNN series methods, or methods based on large language models, have always been limited to capturing the dynamics of the tweet's own propagation structure, and have never considered the impact of tweet structure changes on detection results from the global level of the entire propagation.

[0010] The present application is different from the prior art as follows:

[0011] Technical Comparison with Patent CN202210025558.X, "A Rumor Detection Method and System Based on a Multi-Layer Coding Network"

[0012] Patent CN202210025558.X proposes a rumor detection method and system based on a multi-layer coding network. The method includes the following steps: obtaining all texts to be detected and preprocessing the texts; embedding word slices with a marked vocabulary into the preprocessed texts, converting words in the texts into tokenized words, and then performing vector coding to obtain word vectors corresponding to each text; processing all word vectors to obtain an input vector; inputting the input vector into a pre-trained multi-layer coding network to generate an output vector; processing the output vector to obtain a hidden state vector; and sending the hidden state vector into a hidden layer and a classifier to obtain the probabilities that the text to be detected is detected as each rumor category, and the category with the highest probability is the detection result of the text. Compared with the prior art, this patent not only models the text features of tweets, but also pays attention to temporal information and the interaction relationship between users. By designing a user forwarding network, it realizes the modeling of user relationships across propagation trees. There are essential differences in the implementation methods and technical ideas between the two.

[0013] Technical Comparison with Patent CN119166907A, "A Multi-Modal Rumor Detection Method and System Integrating Multi-Granularity Features"

[0014] Patent CN119166907A proposes a multi-modal rumor detection method and system integrating multi-granularity features. The method includes the following steps: collecting multimedia posts in social media, extracting the text, images, and comments in the posts, and annotating the authenticity labels of the posts to construct a training data set; constructing a multi-modal rumor detection model integrating multi-granularity features, where the multi-modal rumor detection model extracts multi-granularity features in the posts, deeply fuses the text and image features, and realizes the resolution of text-image ambiguity through the text-image similarity weight, and also fully mines the information of comments, utilizes comments from different perspectives, and finally captures the similar information of social media posts of the same category through contrastive learning; training the multi-modal rumor detection model using the training data set; and inputting the text, images, and comments of the multimedia posts to be detected into the trained multi-modal rumor detection model to obtain the authenticity labels of the multimedia posts. This patent is different from multi-modal data and focuses on the collection of text data and user information, which has higher practical significance in practical applications. Secondly, this patent deeply explores the propagation structure of rumors and constructs a user forwarding network based on the original tweet propagation tree to better capture the user relationships across propagation trees. There are essential differences in the invention contents between the two.

[0015] For the problem of early detection of rumors on online social networks, the present invention starts from the perspective of the propagation structure of rumors and designs a new structure information extraction algorithm. To solve the problem of insufficient extraction of propagation features, the present invention proposes a "deep walk spanning tree" graph structure characterization algorithm based on the Prim idea. This method uses the walk method based on the spanning tree to replace traditional methods such as random walk and neighbor aggregation to perceive the local structure. Combining with the tree coding algorithm, each seed node can obtain its coding representation of the surrounding neighbor structure, which can better capture the local propagation structure differences of the subgraph. For the limitation problem of propagation structure modeling, the present invention constructs a user retweet network that combines tweet propagation relationships. This network construction method can better solve the problem that a single propagation tree cannot model the interaction relationship with other information in the information propagation process, and can perform structure modeling on the propagation behaviors of different tweets passing through the same user, so as to better capture the impact of the interaction behavior of tweets from the user's perspective on the information propagation structure. This model not only improves the full excavation and effective utilization of propagation information in rumor detection, but also makes rumor detection more accurate, can timely capture information features in many aspects such as tweet content, forwarding relationship, and user attributes, thus providing strong support for relevant decisions. Summary of the Invention

[0016] To solve the above problems, the present invention proposes a rumor detection method based on information propagation structure. By means of the ordered walk of the spanning tree, it can more systematically encode the contribution of each node to the surrounding neighbor structure, thus more accurately reflecting the local propagation characteristics of the subgraph, rather than the unordered random walk method. At the same time, it effectively integrates multiple propagation paths and user behaviors, reveals the associations formed between different information through common users, thus more comprehensively reflecting the complexity of information propagation, rather than the simple accumulation of local propagation information of a single propagation tree, and effectively improves the accuracy and richness of rumor detection.

[0017] To achieve the above object, the technical solution adopted by the present invention is:

[0018] A rumor detection method based on information propagation structure, comprising the following steps:

[0019] (1) Data analysis module:

[0020] Clean and preprocess the Chinese and English hot event data sets, remove irrelevant information, and use the Jieba word segmentation tool to segment the Chinese text; construct a tweet propagation tree according to the relationship between the original tweet and its retweeted tweets;

[0021] (2) Network construction module:

[0022] According to the tweet propagation tree data obtained in step (1), construct a user retweet network. First, for each propagation tree, replace the tweet node with the user node that posted the tweet, retain the original structure of the propagation tree, and obtain the user retweet subgraph. Then, merge all the user retweet subgraphs. Only retain one user node with the same ID, and only retain one edge in the same direction between two identical users. The node attributes remain unchanged, and the edge attributes are merged to form a complete user retweet network.

[0023] (3) Structure perception module:

[0024] According to the user retweet network obtained in step (2), combined with the retweet attributes such as the number of retweet edges provided by the propagation tree, extract the structure encoding. First, use the WL algorithm-based method to extract the global structure encoding. Then, capture the local structure information from the perspective of user nodes through the "deep walk spanning tree" algorithm, represent it in the form of a tree structure, convert this tree structure into a binary tree and perform local structure encoding. Finally, splice the local structure encoding and the global structure encoding to obtain the complete structure encoding representation.

[0025] (4) Propagation pattern module:

[0026] According to the text content information obtained in step (1), the user retweet network obtained in step (2), and the complete structure encoding obtained in step (3), extract the propagation node representation. First, use the preprocessed text content information and use the pre-trained large language model BERT model to extract text content features to obtain the content encoding. Then, based on the user retweet network, calculate the global, local, and relative position encodings using the Laplacian matrix, intimacy relationship, and shortest path algorithm respectively, and splice the three encodings to obtain the complete position encoding. Finally, splice the content encoding, position encoding, and structure encoding, and input them into the Graph Transformer model to extract the propagation node representation.

[0027] (5) Rumor detection module:

[0028] Input the propagation node representation obtained in step (4) into the multi-layer perceptron MLP model, and then pass through the ReLU activation function and the Softmax regression function to obtain the authenticity label of the tweet.

[0029] As a further improvement of the present invention, after preprocessing and tokenizing the original dataset in step (1), construct a tweet propagation tree, specifically including the steps;

[0030] (1-1) The web links, user mentions and platform tags in the data are converted into a unified format through regular matching, and the text content is filtered based on the stop word list. Then, for Chinese data, the Jieba word segmentation tool is used to segment the Chinese text, filter out the valid information and store it by event, which is represented as the event set E. v ={E1,E2,…,E m}, where E i ={X i ,U i ,T i} represents the information of each event, X i ,U i ,T i They are respectively represented as the text content matrix of event i, the user information set, and the tweet release time set;

[0031] (1-2) For the preprocessed information, we rely on the propagation tree set composed of the original tweet and its forwarded tweets. First, we transform the original tweet x0 and all forwarded tweets R = {x0, x1, …, x n} as nodes in the graph, forming a node set V t ={x0,x1,…,x n}, then, according to the forwarding relationship, establish edge connection p ij =(x i ,x j ), and finally construct the propagation tree G t = {V t ,P t}, where V t , represents the node set, including the original tweet and all forwarded tweets, P t represents the edge set, edge p ij ∈P t Indicates that the i-th user forwarded the j-th tweet, adds the propagation tree to the event set, and finally obtains the tweet event data E i = {G i ,X i ,U i ,T i}.

[0032] As a further improvement of the present invention, step (2) specifically comprises:

[0033] (2-1) To mine the relationship between multiple propagation trees, a user forwarding network that can unify multiple propagation trees is constructed. According to the event set E obtained in step (1), v ={E1,E2,…,E m}, where E i = {G i ,X i,U i ,T i}, where G, X, U, and T are the propagation tree, text content matrix, user information set, and tweet release time set respectively. For each propagation tree G i = {V i , P i}, where V i is the node set and P i is the edge set. Construct the user retweet subgraph G′ i = {V′ i , P′ i}. Traverse each node j of each propagation tree G i , extract the attribute u of the publishing user to which the tweet node x j belongs, replace the tweet node x j with the user node u j that publishes the tweet, and form the user node set V′ j = {v0.u, v1.u, …, v i .u}. At the same time, retain the tweet forwarding relationship in the original propagation tree, replace the tweet node p n = (x i , x ij ) in the edge set P i with the user node p j = (v ij .u, v i .u), and form the user edge set P′ j . Thus, form the user retweet subgraph G i ′ = {V i ′, P i ′}, and finally obtain the user retweet subgraph set G′ = {G′1, G′2, …, G i ′};

[0034] (2-2) After constructing the user retweet subgraph set G′ = {G′1, G′2, …, G m ′}, merge all user retweet subgraphs. In the first step, process the nodes in the user retweet subgraphs to construct the node set V′. Traverse each user retweet subgraph G′ m in turn. For each node v i , check whether the user ID of this node exists in V′. If it does not exist, create a new node v′ i in V′, initialize the attribute lists x' id and t′ id , and then add the text content x id and time t i of the current node to the corresponding node v′ i . idAttribute list x' id and t' id in;

[0035] Second, process the edges in the user retweet subgraph, construct the edge set E', and sequentially traverse each user retweet subgraph G' i , for each edge p k =(v x ,v y ), obtain the start node v k and the end node v x of the edge p y , find the user nodes corresponding to v x and v y in V', which are respectively represented as v' x ,v' y , check whether the edge e' k =(v' x ,v' y ) exists in E'. If not, create the attribute list of e′ k , add the attributes of e k to the list, initialize the num_edge attribute of e' k to 1. If it exists, add the attributes of ek to the attribute list of e' k , and increment the numerical value of the num_edge attribute of e′ k by 1; Finally, obtain the complete user retweet network U = {V', E'}.

[0036] As a further improvement of the present invention, step (3) specifically includes:

[0037] (3-1) According to the user forwarding network U = {V', E'} obtained in step (2), combined with attributes such as the number of forwarding edges provided by the propagation tree, extract the structure encoding. First, use the WL algorithm-based method to extract the global structure encoding of the user forwarding network. The specific steps are as follows:

[0038] First, for each node v'∈V' in the graph, initialize the node label. Use the Hash hash function to color according to its node attributes, and encode the original attributes of the node into an initial color value, which is expressed as:

[0039]

[0040] where, represents the label of node v' at iteration 0. After that, start the iteration. For each node v′, calculate the label set N(v') of its neighbor nodes, that is, the current label set of all nodes directly connected to v'. Then update the node label and remove duplicates. Use the Hash function or other feature mapping function f to combine the current label of node v and the set of labels N(v') of its neighboring nodes:

[0041]

[0042] where Neigh U (v′) represents the set of neighboring nodes of the node on graph U. To avoid label inflation, it is necessary to ensure that for two nodes with the same label, their newly mapped labels are also the same:

[0043] and N(v′) = N(u')

[0044] When the iterative process meets any of the following termination conditions:

[0045] First, when the node labels do not change in two consecutive iterations, that is, for all v′ ∈ V′, there is:

[0046]

[0047] Second, when the number of iterations l reaches the preset maximum value L max ;

[0048] Finally, when the iteration stops, the labels of all nodes have converged. The final labels are combined into a global structure feature encoding in matrix form, denoted as SE Global :

[0049]

[0050] (3 - 2) In the second step, the "depth - first search spanning tree" algorithm is used to capture the local structure information from the perspective of user nodes and represent it in a tree structure. The tree structure representation method based on the traversal algorithm is used to serialize the tree structure. Finally, the Skip - gram model is used to embed the sequence to obtain the local structure encoding. The specific steps are as follows:

[0051] First, according to the user forwarding network U = {V′, E'} obtained in step (2), for each node v' ∈ V′ to be represented in the graph, create an empty tree Tree v′ to record the results of the traversal. Create an empty set of reachable nodes S to record the nodes that have been traversed and an empty set of reached nodes A to record the nodes that can be reached in the next step during the traversal. Initialize the set of reached nodes A = {v′}, indicating the visited nodes, and initialize the edge set to store the edge information in the tree;

[0052] For each node v′, perform ST num independent depth - first searches. Each search visits at most ST lenA new node, the starting point of the walk is set to start from the current node u = v', and the step counter t = 0 is initialized;

[0053] Then start traversing. In each walk, starting from the current node u, calculate the one-hop reachable nodes of all nodes in A on G as S', calculate the transition probabilities of the nodes in S according to the walk parameters p and q, and sample a node w ∈ Neigh U (u), where Neigh U (u) is the neighbor set of node u, and add the node w to the set A:

[0054] A = A ∪ {w}

[0055] At the same time, generate an edge e = (u, w), and add it to the tree Tree according to the edge e = (u, w); when ST len new nodes have been visited, or when unable to move forward, terminate the current walk, and finally obtain the t-th spanning tree Tree of node v' v′ ;

[0056] When the relative order between the child nodes of each node in a multi-way tree is fixed, the tree is uniquely determined as a binary tree, and the process is reversible. Therefore, after obtaining ST num multi-way trees containing ST len nodes Convert the multi-way tree to a binary tree And perform pre-order traversal and in-order traversal on the binary tree to obtain an encoded sequence representation that concatenates the results of pre-order traversal and in-order traversal, PreList and MidList with a length of 2ST len ;

[0057] Subsequently, use the Skip-gram model to embed the structural features of the encoded sequence representation L, and extract the local structural encoding SE of the node Local ;

[0058] (3-3) Step 3, concatenate the global structural encoding SE Global and the local structural encoding SE Local to obtain the complete structural encoding SE:

[0059]

[0060] where represents the concatenation of feature vectors.

[0061] As a further improvement of the present invention, step (4) specifically includes:

[0062] (4-1) For the text content information obtained in step (1), for the tweet set X = {x1, x2, …, x n}, where n represents the number of tweets in the set, and for any a < b, the publication time of tweet x a is no later than that of tweet x b . x i is also represented as [w i,1 , w i,2 , …, w i,m , where w i,j represents the j-th word in tweet x i , and m is the number of words in tweet x i . First, the tweets are divided at the sentence level and then fed into the pre-trained BERT model to obtain their word vector representations, and each sentence corresponds to a word vector matrix. Then, the word-level word vectors are fed into the Bi-LSTM model to learn their context relationships. For word w i,j , in the forward LSTM, the semantic information carried by the previous word w i,j-1 is passed to it. After learning the semantics of w i,j-1 and the previous words, some unimportant information is forgotten and then passed to the next word w i,j+1 . In the reverse LSTM, the information of the word sequence of w i,j and the subsequent words that w i,j+1 did not obtain in the forward process is also passed to it with attenuation through w i,j+1 , so that the Bi-LSTM can retain some global dependency information while paying attention to local sequence information. Its formal representation is:

[0063]

[0064] where, and are the vector representations of word w i,j+1 in the forward and reverse LSTM hidden layers respectively. The final representation h i,j of word w i,j+1 is the concatenation of the forward representation and the reverse representation . Then, it is input into the Attention layer to model the importance of each word in the sentence, calculate its attention coefficient, and obtain the vector representation of the entire sentence by weighted summation with the content vector. For the hidden vector h i,j of the (a - 1)-th layer that has been calculated, the calculation method of the attention coefficient e i,j at the position of word w i,j in the next layer is:

[0065] ei,j = Attention(s a-1 ,h i,j )

[0066] After that, the obtained attention coefficients are normalized through the Softmax layer, and the attention weight a i,j of the word w i,j :

[0067]

[0068] The final representation c i of the tweet x is calculated by weighted summation, i.e.: i :

[0069]

[0070] Finally, the text content encoding matrix CE = C = {c1, c2, …, c n};

[0071] (4 - 2) According to the user forwarding network U = {V′, E′} obtained in step (2), first, the user forwarding network is converted into the corresponding complete Laplacian matrix to obtain the global position encoding PE Global , and the formula is as follows:

[0072] PE Global = I - D -1 / 2 AD 1 / 2 = Y T ΛY

[0073] where A is the complete network adjacency matrix, D is the complete network degree matrix, and Λ and Y correspond to the eigenvalues and eigenvectors respectively;

[0074] Then, since the user retweet network is formed by fusing the propagation subgraphs converted from the original propagation tree, the original propagation tree naturally forms a local clustering. Therefore, this local clustering, i.e., the user retweet network, etc., is used for the intimacy ranking, i.e.:

[0075] PE Local = P(v)

[0076] where P(v) represents that for the target node v, other nodes are sorted in descending order according to the intimacy scores calculated with v to obtain the intimacy distance between two nodes, i.e., the intimacy matrix, and thus the local position encoding PE Local is obtained;

[0077] Subsequently, the tweet release time attribute is used to judge the propagation direction relationship between nodes, and the relative position encoding PE Relative is calculated through the Floyd shortest path algorithm, i.e.:

[0078] PE Relative = Floyd(A)

[0079] Among them, Floyd(A) is the node shortest distance matrix obtained by performing the Floyd algorithm on the complete network adjacency matrix A;

[0080] Finally, concatenate the global position encoding PE Global , the local position encoding PE Local , and the relative position encoding PE Relative to obtain the complete position encoding PE:

[0081]

[0082] Among them represents the concatenation of feature vectors;

[0083] (4-3) According to the structure encoding SE obtained in step (3), the text content encoding CE obtained in step (4-1), and the position encoding PE obtained in step (4-2), represent the content encoding of each tweet node in the text content encoding as Map the features of nodes and edges to d dimensions through linear transformation. The specific calculation formula is:

[0084]

[0085] Among them, is the feature of node i, is the feature of the edge between node i and node j, A 0 and B 0 are projection matrices, a 0 and b 0 are the bias parameters of the linear mapping;

[0086] Similarly, perform linear mapping on the position and structure encodings to the same dimension, and add them to the node features:

[0087]

[0088] Among them, is the feature encoding of the position encoding matrix, is the feature encoding of the structure encoding matrix, C 0 and D 0 are projection matrices, c 0 and d 0 are the bias parameters of the linear mapping, is the final initial feature representation of the node;

[0089] Input the final initial node representation sequence into the Graph Transformer network, and take the hidden state after the last iteration as the propagated node representation. Each layer of the Graph Transformer consists of the following parts: self-attention mechanism, feature aggregation, normalization, and residual connection. For the node update of a certain layer, the specific formula is as follows:

[0090]

[0091] Among them, || represents concatenation, represents the attention weight, and the specific calculation formula is:

[0092]

[0093] After that, pass the output result to a single-layer feed-forward neural network, and calculate the residual connection and normalization to obtain the embedded representation of the node:

[0094]

[0095] For the output of the last layer We denote it as the propagated node representation Z, that is:

[0096]

[0097] As a further improvement of the present invention, step (5) specifically includes:

[0098] According to the propagated node representation Z obtained in step (4), use a multi-layer perceptron MLP to calculate the authenticity score of the tweet. The specific calculation formula is:

[0099] y t = softmax(W Y (σ(W z ·Z)))

[0100] Among them, W Y and W Z are the weight matrices of the MLP, Z is the final node representation, σ is the ReLU activation function. For y t , the classification label corresponding to the item with the highest authenticity score is the authenticity label of the predicted information.

[0101] Beneficial effects: Compared with the prior art, the present invention adopts the above technical solutions and has the following advantages:

[0102] (1)Regarding the problem of propagation structure differences, the present invention designs a local propagation feature modeling method based on depth-first random walk on a spanning tree. By systematically encoding the contribution of each node to the surrounding neighbor structure through an ordered random walk on the spanning tree, the local propagation characteristics of the subgraph can be more accurately reflected, rather than using an unordered random walk method. This method can better capture the local propagation structure differences of the subgraph and greatly improve the accuracy of rumor detection;

[0103] (2)Regarding the limitation problem of information interaction modeling, the present invention constructs a user retweet network that combines tweet propagation relationships. By integrating multiple propagation paths and user behaviors, the associations formed between different information through common users are revealed, so as to more comprehensively reflect the complexity of information propagation, rather than simply adding up the local propagation information of a single propagation tree. This method can effectively improve the full exploration and effective utilization of propagated information, making rumor detection more comprehensive and accurate;

[0104] (3)Regarding the problem of rumor detection effect, the present invention proposes a rumor detection method based on information propagation structure. By systematically encoding node neighbor relationships and integrating multi-dimensional information features, information such as tweet content, retweet relationships, and user attributes is comprehensively modeled, so as to more accurately express the information propagation law and user behavior characteristics. This method can timely capture multi-faceted information features, provide more comprehensive data support for rumor detection, and make relevant decisions more scientific and reliable. Brief Description of the Drawings

[0105] Figure 1 is the overall framework diagram of the present invention;

[0106] Figure 2 is a schematic diagram of the tweet propagation tree structure;

[0107] Figure 3 is a schematic diagram of the network construction module;

[0108] Figure 4 is a schematic diagram of global structure encoding;

[0109] Figure 5 is a schematic diagram of the depth-first random walk algorithm;

[0110] Figure 6 is a schematic diagram of local structure encoding;

[0111] Figure 7 is a schematic diagram of text content encoding;

[0112] Figure 8 is a schematic diagram of the propagation node representation based on the Graph Transformer network. Detailed Embodiment

[0113] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0114] As Figure 1 shown, a rumor detection method based on an information dissemination structure of the present invention includes the following steps:

[0115] 1. Data analysis module

[0116] Clean and preprocess the original datasets PHEME and Misinfdect, remove irrelevant information, and use the Jieba word segmentation tool to perform word segmentation on Chinese texts; construct a tweet propagation tree according to the relationship between the original tweet and its retweeted tweets;

[0117] (1) For the web links, user mentions, and platform tags in the data, convert them into a unified format by regular matching, filter the text content based on the stop word list, and then, for Chinese data, use the Jieba word segmentation tool to perform word segmentation on Chinese texts, screen out the valid information and store it according to events, denoted as event set E v ={E1, E2, …, E m}}, where E i ={X i , U i , T i} represents the information of each event, and X i , U i , T i represent the text content matrix, user information set, and tweet release time set of event i, respectively;

[0118] (2) For the preprocessed information, relying on the propagation tree set composed of the original tweet and its retweeted tweets, first, use the original tweet x0 and all retweeted tweets R = {x0, x1, …, x n} as nodes in the graph to form node set V t ={x0, x1, …, x n}, then, according to the retweet relationship, establish edge connection p ij =(x i , x j ), and finally construct the propagation tree as G t ={V t , P t}, where V t represents the node set, including the original tweet and all retweeted tweets, and P t represents the edge set, and edge p ij ∈ P tIndicates that the i-th user has retweeted the j-th tweet. The propagation tree structure is as follows Figure 2 shown. Add the propagation tree to the event set, and finally obtain the tweet event data E i ={G i , X i , U i , T i}}.

[0119] 2. Network Construction Module

[0120] According to the tweet propagation tree data obtained in step (1), construct a user retweet network. First, for each propagation tree, replace the tweet node with the user node that posted the tweet, retain the original structure of the propagation tree, and obtain the user retweet subgraph; then, merge all user retweet subgraphs, only retain one user node with the same id, only retain one same-direction edge between two identical users, keep the node attributes unchanged, and merge the edge attributes to form a complete user retweet network. The network construction module is as follows Figure 3 shown.

[0121] (1) To mine the relationships between multiple propagation trees and solve the problem that the traditional propagation structure cannot represent the same user in different propagation trees, it is necessary to construct a user retweet network that can unify multiple propagation trees. According to the event set E v ={E1, E2, …, E m}, where E i ={G i , X i , U i , T i}, G, X, U, and T are the propagation tree, text content matrix, user information set, and tweet release time set respectively. For each propagation tree G i ={V i , P i}, where V i is the node set and P i is the edge set. Construct a user retweet subgraph G′ i ={V′ i , P′ i}. Traverse each node j of each propagation tree G i , extract the attribute u j of the user who posted the tweet node x j to which it belongs, replace the tweet node x j with the user node u j that posted the tweet, form the user node set V′ i ={v0.u, v1.u, …, v n .u}, and at the same time retain the tweet retweet relationship in the original propagation tree, and use the edge set Pi The tweet node p in ij =(x i , x j ) is replaced with the user node p who posted the tweet ij =(v i .u, v j .u), forming the user edge set P' i , thus forming the user retweet subgraph G i '={V i ', P i '}, and finally obtaining the user retweet subgraph set G'={G'1, G'2,..., G m '}.

[0122] (2) After constructing the user retweet subgraph set G'={G'1, G'2,..., G m '}, all user retweet subgraphs are merged. In the first step, the nodes in the user retweet subgraph are processed to construct the node set V'. Each user retweet subgraph G' is traversed in turn i . For each node v i , check whether the user ID of this node exists in V'. If it does not exist, create a new node v' in V id , initialize the attribute lists x' id and t' id , and then add the text content x i and time t i of the current node to the corresponding node v' id 's attribute lists x' id and t' id ;

[0123] In the second step, the edges in the user retweet subgraph are processed to construct the edge set E'. Each user retweet subgraph G' is traversed in turn i . For each edge p k =(v x , v y ), obtain the start node v k and the end node v x of the edge p y . Find the user nodes corresponding to v x and v y in V', which are respectively represented as v' x , v' y . Check whether the edge e' k =(v′ x , v' y ) exists in E'. If it does not exist, create the attribute list of e′ k , add the attributes of e k to the list, and initialize e'k The num_edge attribute of is 1. If it exists, then e k 's attributes are added to e'. k into the attribute list of, and e' k 's num_edge attribute value is incremented by 1; Finally, the complete user retweet network U = {V', E'} is obtained.

[0124] 3. Structure Perception Module

[0125] Based on the user retweet network obtained in step (2), combined with retweet attributes such as the number of retweet edges provided by the propagation tree, extract the structure encoding. First, use the WL algorithm-based method to extract the global structure encoding; then, capture the local structure information from the perspective of user nodes through the "DeepWalk spanning tree" algorithm and represent it in the form of a tree structure, and convert this tree structure into a binary tree and perform local structure encoding; finally, splice the local structure encoding and the global structure encoding to obtain the complete structure encoding representation;

[0126] (1) According to the user retweet network U = {V', E'} obtained in step (2), combined with attributes such as the number of retweet edges provided by the propagation tree, extract the structure encoding. In the first step, use the WL (Weisfeiler-Lehman test) algorithm-based method to extract the global structure encoding of the user retweet network, as Figure 4 shown. The specific steps are as follows:

[0127] First, for each node v' ∈ V' in the graph, initialize the node label. Use the Hash hash function to color according to its node attributes, and encode the original attributes of the node into an initial color value, expressed as:

[0128]

[0129] where represents the label of node v' at iteration 0. After that, start the iteration. For each node v', calculate the label set N(v') of its neighbor nodes, that is, the current label set of all nodes directly connected to v'. Then update the node label and remove duplicates, and use the Hash function or other feature mapping function f to combine the current label of node v and the label set N(v') of its neighbor nodes:

[0130]

[0131] where Neigh U (v') represents the set of neighbor nodes of the node on graph U. To avoid label inflation, it is necessary to ensure that for two nodes with the same label, their newly mapped labels are also the same:

[0132]

[0133] When the iterative process satisfies any of the following termination conditions:

[0134] First, when the node labels do not change in two consecutive iterations, that is, for all v′ ∈ V′, there is:

[0135]

[0136] Second, when the number of iterations l reaches the preset maximum value L max ;

[0137] Finally, when the iteration stops, the labels of all nodes have converged. The final labels are combined into a global structure feature encoding in matrix form, denoted as SE Global :

[0138]

[0139] (2) In the second step, the "depth-first search spanning tree" algorithm is used to capture the local structure information from the perspective of user nodes and represent it in a tree structure. The tree structure representation method based on the traversal algorithm is used to serialize the tree structure, and finally the Skip-gram model is used to embed the sequence to obtain the local structure encoding. The specific steps are as follows:

[0140] First, according to the user forwarding network U = {V′, E'} obtained in step (2), for each node v' ∈ V′ to be represented in the graph, create an empty tree Tree v′ to record the traversal results, create an empty set S of reachable nodes to record the nodes that have been traversed and an empty set A of reached nodes to record the nodes that can be reached in the next step during the traversal. Initialize the set A of reached nodes = {v′}, indicating the nodes that have been visited, and initialize the edge set to store the edge information in the tree;

[0141] Perform ST num independent depth-first searches for each node v′. Each search visits at most ST len new nodes. The starting point of the search is set to the current node u = v', and initialize the step counter t = 0;

[0142] Then start the traversal. In each search, starting from the current node u, calculate the one-hop reachable nodes of all nodes in A on G as S’. Calculate the transition probabilities of the nodes in S according to the traversal parameters p and q, and sample a node w ∈ Neigh U (u), where Neigh U (u) is the neighbor set of node u, and add the node w to the set A:

[0143] A = A ∪ {w} #(6)

[0144] Meanwhile, generate an edge e=(u, w), and add it to the tree Tree according to the edge e=(u, w); when ST len new nodes have been visited, or when it is impossible to move forward, terminate the current walk, and finally obtain the t-th spanning tree Tree v′ of the node v'. The "depth walk spanning tree" algorithm is as Figure 5 shown;

[0145] When the relative order between the children of each node in a multi-way tree is fixed, the tree can be uniquely represented as a binary tree, and the process is reversible. Therefore, after obtaining ST num multi-way trees containing ST len nodes Convert the multi-way tree to a binary tree using the "left child, right sibling" method and perform pre-order traversal and in-order traversal on the binary tree to obtain the encoded sequence representation that concatenates the results of pre-order traversal and in-order traversal, PreList and MidList, into a tree structure with a length of 2ST len ;

[0146] Subsequently, use the Skip-gram model to embed the structural features of the encoded sequence representation L, and extract the local structural encoding SE Local of the node, as Figure 6 shown;

[0147] (3) In the third step, concatenate the global structural encoding SE Global and the local structural encoding SE Local to obtain the complete structural encoding SE:

[0148]

[0149] where represents the concatenation of feature vectors.

[0150] 4. Propagation Mode Module

[0151] According to the text content information obtained in step (1), the user forwarding network obtained in step (2), and the complete structure encoding obtained in step (3), the propagation node representation is extracted. First, using the preprocessed text content information, the pre-trained large language model BERT model is used to extract text content features to obtain content encoding; then, based on the user forwarding network, the Laplacian matrix, intimacy relationship, and shortest path algorithm are used to calculate the global, local, and relative position encodings respectively, and the three encodings are concatenated to obtain the complete position encoding; finally, the content encoding, position encoding, and structure encoding are concatenated and input into the Graph Transformer model to extract the propagation node representation;

[0152] (1) For the text content information obtained in step (1) for the tweet set X = {x1, x2, …, x n}, where n represents the number of tweets in the set, and for any a < b, the publication time of tweet x a is not later than that of tweet x b , x i can also be expressed as [w i,1 , w i,2 , …, w i,m , where w i,j represents the j-th word in tweet x i , m is the number of words in tweet x i . First, the tweets are divided at the sentence level and then fed into the pre-trained BERT model to obtain their word vector representations, and each sentence corresponds to a word vector matrix; then, the word-level word vectors are fed into the Bi-LSTM model to learn their context relationships. Since Bi-LSTM performs the LSTM process in both the forward and reverse orders, it can learn the connections between words in both the forward and reverse directions while retaining its long-distance dependence characteristics. For the word w i,j , in the forward LSTM, the semantic information carried by the previous word w i,j-1 is passed to it. After learning the semantics of w i,j-1 and the previous words, some unimportant information is forgotten and then passed to the next word w i,j+1 . In the reverse LSTM, the information of the word sequence of w i,j that w i,j+1 did not obtain in the forward process and the subsequent words is also passed to it with attenuation through w i,j+1 , so that Bi-LSTM can retain some global dependence information while focusing on local sequence information. Its formal representation is:

[0153]

[0154] Among them, and are the vector representations of the word w i,j+1 in the forward and backward LSTM hidden layers respectively. The final representation h i,j+1 of the word w i,j is the forward representation concatenated with the backward representation ; then, it is input into the Attention layer to model the importance of each word in the sentence, calculate its attention coefficient, and obtain the vector representation of the entire sentence by weighted summation with the content vector. For the hidden vector h i,j calculated at the (a - 1)-th layer, the attention coefficient e i,j at the position of the word w i,j in the next layer is calculated as follows:

[0155] e i,j = Attention(s a-1 , h i,j )#(11)

[0156] After that, the obtained attention coefficients are normalized through the Softmax layer to calculate the attention weight a i,j of the word w i,j :

[0157]

[0158] The final representation c i of the tweet x i is calculated by weighted summation:

[0159]

[0160] Finally, the text content encoding matrix CE = C = {c1, c2,..., c n} is obtained, and the content encoding is as shown in Figure 7 .

[0161] (2) According to the user forwarding network U = {V′, E'} obtained in step (2), first, the user forwarding network is converted into the corresponding complete Laplacian matrix to obtain the global position encoding PE Global , and the formula is as follows:

[0162]

[0163] where A is the complete network adjacency matrix, D is the degree matrix of the complete network, and Λ and Y correspond to the eigenvalues and eigenvectors respectively.

[0164] Then, since the user retweet network is formed by fusing propagation subgraphs converted from the original propagation tree, a local clustering is naturally formed in the original propagation tree. Therefore, this local clustering, i.e., the user retweet network, etc., is used for intimacy ranking, that is:

[0165] PE Local = P(v)#(15)

[0166] Among them, P(v) represents the intimacy distance between two nodes obtained by sorting other nodes in descending order of their intimacy scores calculated with v for the target node v, that is, the intimacy matrix. From this, the local position encoding PE is obtained. Local .

[0167] Subsequently, the release time attribute of the tweet is used to judge the propagation direction relationship between nodes, and the relative position encoding PE is calculated through the Floyd shortest path algorithm. Relative , that is:

[0168] PE Relative = Floyd(A)#(16)

[0169] Among them, Floyd(A) is the shortest distance matrix of nodes calculated by executing the Floyd algorithm on the complete network adjacency matrix A.

[0170] Finally, the global position encoding PE Global , the local position encoding PE Local , and the relative position encoding are concatenated to obtain the complete position encoding PE: Relative Among them

[0171]

[0172] where represents the concatenation of feature vectors.

[0173] (3) According to the structure encoding SE obtained in step (3), the text content encoding CE obtained in step (4), and the position encoding PE, the content encoding of each tweet node in the text content encoding is expressed as The features of the node and the edge are mapped to d dimensions through a linear transformation. The specific calculation formula is:

[0174]

[0175] Among them, is the feature of node i, is the feature of the edge between node i and node j, A 0 and B 0 are projection matrices, a 0 and b 0 are the bias parameters of the linear mapping;

[0176] Similarly, perform a linear mapping of the position and structure encodings to the same dimension and add them to the node features:

[0177]

[0178] where is the feature encoding of the position encoding matrix, is the feature encoding of the structure encoding matrix, C 0 and D 0 are projection matrices, c 0 and d 0 are the bias parameters of the linear mapping, is the final initial node feature representation;

[0179] Input the final initial node representation sequence into the Graph Transformer network, as shown in Figure 8 . Take the hidden state after the last iteration as the propagated node representation. Each layer of the Graph Transformer mainly consists of the following parts: self-attention mechanism, feature aggregation, normalization, and residual connection. For the node update of a certain layer, the specific formula is:

[0180]

[0181] where || represents concatenation, represents the attention weight, and the specific calculation formula is:

[0182]

[0183] After that, pass the output result to a single-layer feedforward neural network (Feedforward Neural Network, FNN), and calculate the residual connection and normalization to obtain the embedded representation of the node:

[0184]

[0185] For the output of the last layer we take it as the propagated node representation Z, that is:

[0186]

[0187] 5. Rumor Detection Module

[0188] According to the propagated node representation Z obtained in step (4), use a multi-layer perceptron (MLP) to calculate the authenticity score of the tweet. The specific calculation formula is:

[0189] y t = softmax(WY (σ(W z ·Z)))#(28)

[0190] Among them, W Y and W Z are the weight matrices of the MLP, Z is the representation of the last node, and σ is the ReLU activation function. For y t , the classification label corresponding to the item with the highest authenticity score is the authenticity label of the predicted information.

[0191] The above is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in any other form. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.

Claims

1. A rumor detection method based on information propagation structure, characterized by: The following steps are involved: (1) Data analysis module: Clean and preprocess the Chinese and English hot event data sets to remove irrelevant information, and use the Jieba word segmentation tool to segment the Chinese text; construct a tweet propagation tree based on the relationship between the original tweet and its forwarded tweets; (2) Network building blocks: Based on the tweet propagation tree data obtained in step (1), a user forwarding network is constructed. First, for each propagation tree, the tweet node is replaced with the user node that posted the tweet, the original structure of the propagation tree is retained, and the user forwarding subgraph is obtained. Then, all user forwarding subgraphs are merged, only one user node with the same ID is retained, and only one same-direction edge between two identical users is retained. The node attributes remain unchanged, and the edge attributes are merged, thereby forming a complete user forwarding network. (3) Structure perception module: According to the user forwarding network obtained in step (2), combined with the forwarding attributes such as the number of forwarding edges provided by the propagation tree, the structural code is extracted. First, the global structural code is extracted using the WL algorithm; then, the local structural information from the user node's perspective is captured by the "deep walk spanning tree" algorithm and represented in the form of a tree structure, and the tree structure is converted into a binary tree and local structural coding is performed; finally, the local structural code and the global structural code are spliced ​​to obtain a complete structural code representation; (4) Propagation mode module: According to the text content information obtained in step (1), the user forwarding network obtained in step (2) and the complete structure code obtained in step (3), the propagation node representation is extracted. First, the pre-processed text content information is used to extract the text content features using the pre-trained large language model BERT model to obtain the content code; then, based on the user forwarding network, the global, local and relative position codes are calculated using the Laplacian matrix, the intimacy relationship and the shortest path algorithm respectively, and the three codes are spliced ​​to obtain the complete position code; finally, the content code, the position code and the structure code are spliced ​​and input into the GraphTransformer model to extract the propagation node representation; (5) Rumor Detection Module: The propagation node representation obtained in step (4) is input into the multi-layer perceptron MLP model, and then passes through the ReLU activation function and the Softmax regression function to obtain the authenticity label of the tweet.

2. The rumor detection method based on information propagation structure according to claim 1 is characterized by: After preprocessing and word segmentation of the original data set in step (1), a tweet propagation tree is constructed, which specifically includes the following steps: (1-1) The web links, user mentions and platform tags in the data are converted into a unified format through regular matching, and the text content is filtered based on the stop word list. Then, for Chinese data, the Jieba word segmentation tool is used to segment the Chinese text, filter out the valid information and store it by event, which is represented as the event set E. v ={E1,E2,…,E m }, where E i ={X i ,U i ,T i } represents the information of each event, X i ,U i ,T i They are respectively represented as the text content matrix of event i, the user information set, and the tweet release time set; (1-2) For the preprocessed information, we rely on the propagation tree set composed of the original tweet and its forwarded tweets. First, we transform the original tweet x0 and all forwarded tweets R = {x0, x1, …, x n } as nodes in the graph, forming a node set V t ={x0,x1,…,x n }, then, according to the forwarding relationship, establish edge connection p ij =(x i ,x j ), and finally construct the propagation tree G t = {V t ,P t }, where V t , represents the node set, including the original tweet and all forwarded tweets, P t represents the edge set, edge p ij ∈P t Indicates that the i-th user forwarded the j-th tweet, adds the propagation tree to the event set, and finally obtains the tweet event data E i = {G i ,X i ,U i ,T i }.

3. The rumor detection method based on information propagation structure according to claim 1 is characterized in that: Step (2) specifically includes: (2-1) To mine the relationship between multiple propagation trees, a user forwarding network that can unify multiple propagation trees is constructed. According to the event set E obtained in step (1), v ={E1,E2,…,E m }, where E i = {G i ,X i ,U i ,T i }, G, X, U, T are propagation trees, text content matrices, user information sets, and tweet release time sets, respectively. For each propagation tree G i = {V i ,P i }, where V i is a set of nodes, P i As the edge set, construct the user retweet subgraph G′ i = {V′ i ,P′ i }, traverse each propagation tree G i For each node j, extract its tweet node x j The publishing user attribute u j , push node x j Replaced with the user node u that posted the tweet j , forming the user node set V′ i ={v0.u,v1.u,…,v n .u}, while retaining the tweet forwarding relationship in the original propagation tree, and transforming the edge set P i The tweet node p in ij =(x i ,x j ) is replaced by the user node p that posted the tweet ij =(v i .u,v j .u), forming the user edge set P′ i , thus forming the user retweet subgraph G i ′={V i ′,P i ′}, and finally we get the user retweet subgraph set G′={G′1,G′2,…,G m '}; (2-2) Construct the user retweet subgraph set G′={G′1,G′2,…,G m ′}, merge all user retweet subgraphs. In the first step, process the nodes in the user retweet subgraph, construct the node set V′, and traverse each user retweet subgraph G′ in turn. i , for each node v i , check whether the user ID of the node exists in V′, if not, create a new node v′ in V′ id , initialize the attribute list x′ of the node id and t′ id , and then the text content of the current node x i and time t i Add to the corresponding node v′ id The property list of x' id and t' id middle; The second step is to process the edges in the user retweet subgraph, construct the edge set E′, and traverse each user retweet subgraph G′ in turn. i , for each edge p k =(v x ,v y ), get edge p k The starting node v x and the terminal node v y , find the value corresponding to v in V′ x and v y The user nodes are represented as v′ x ,v′ y , check edge e' k =(v′ x ,v′ y ), whether it exists in E', if not, create e' k The property list of e k Add the properties of the list and initialize e' k The num_edge attribute of is 1, if it exists, e k Add the properties of e′ k in the attribute list and replace e' k The num_edge attribute value of is added by 1; finally, the complete user retweet network U = {V′, E′} is obtained.

4. The rumor detection method based on information propagation structure according to claim 1 is characterized by: Step (3) specifically includes: (3-1) According to the user forwarding network U = {V′, E′} obtained in step (2), combined with the attributes such as the number of forwarding edges provided by the propagation tree, the structural code is extracted. In the first step, the global structural code of the user forwarding network is extracted using the WL algorithm. The specific steps are as follows: First, for each node v'∈V' in the graph, initialize the node label, use the Hash function to color it according to its node attributes, and encode the original attributes of the node into an initial color value, expressed as: in, represents the label of node v' at iteration 0. After that, the iteration starts. For each node v', the label set N(v') of its neighbor nodes is calculated, that is, the current label set of all nodes directly connected to v'. Then the node labels are updated and duplicated. The hash function or other feature mapping function f is used to combine the current label of node v And the label set N(v') of its neighbor nodes: Among them, Neigh U (v′) represents the set of neighbor nodes of the node on the graph U. In order to avoid label expansion, it is necessary to ensure that for two nodes with the same label, their new labels after mapping are also the same: And N(v') = N(u') The iteration process terminates when any of the following conditions are met: First, when the node label does not change in two consecutive iterations, that is, for all v'∈V': Second, when the number of iterations l reaches the preset maximum value L max ; Finally, when the iteration stops, the labels of all nodes have converged, and the final labels are combined into a global structural feature encoding in matrix form, denoted as SE Global : (3-2) The second step is to capture the local structure information from the user node perspective through the "deep walk spanning tree" algorithm and represent it in a tree structure. The tree structure representation method based on the traversal algorithm is used to serialize the tree structure. Finally, the Skip-gram model is used to embed the sequence to obtain the local structure encoding. The specific steps are as follows: First, according to the user forwarding network U = {V′, E′} obtained in step (2), for each node v′∈V′ to be represented on the way, create an empty tree Tree v′ Used to record the results of the walk, create an empty reachable node set S to record the nodes that have been walked and an empty reached node set A to record the nodes that can be reached in the next step of the walk, initialize the reached node set A = {v′}, indicating the nodes that have been visited, and initialize the edge set Used to store edge information in the tree; For each node v', perform ST num Independent deep walks, each walk visits at most ST len A new node is created, the starting point of the walk is set from the current node u=v′, and the step counter t=0 is initialized; Then start the traversal. In each walk, starting from the current node u, calculate the one-hop reachable node S' of all nodes in A on G, calculate the transition probability of the node in S according to the walk parameters p and q, and sample the node w∈Neigh according to the transition probability distribution U (u), where Neigh U (u) is the neighbor set of node u, add node w to set A: A=A∪{w} At the same time, an edge e = (u, w) is generated and added to the tree Tree according to the edge e = (u, w); when ST has been visited len When a new node is reached or it is impossible to move forward, the current walk is terminated, and finally the t-th spanning tree Tree of node v′ is obtained. v′ ; When the relative order of the child nodes of each node in a multi-branch tree is fixed, the tree can be uniquely represented as a binary tree, and the process is reversible. Therefore, when obtaining ST num Contains ST len Multi-node tree Finally, convert the multi-fork tree into a binary tree Perform pre-order traversal and in-order traversal on the binary tree to obtain the encoding sequence representation of the tree structure by splicing the results of the pre-order traversal and in-order traversal PreList and MidList Length is 2ST len ; Then, the Skip-gram model is used to embed the structural features of the coding sequence representation L, and the local structure encoding SE of the extracted node is Local ; (3-3) The third step is to encode the global structure SE Global and local structure encoding SE Local The complete structural code SE is obtained by splicing: SE=SE blobal ⊕SE Local Where ⊕ represents the concatenation of feature vectors.

5. The rumor detection method based on information propagation structure according to claim 1 is characterized by: Step (4) specifically includes: (4-1) For the text content information obtained in step (1), for the tweet set X = {x1, x2, …, x n}, where n represents the number of tweets in the set, and for any a < b, the posting time of tweet x a is no later than that of tweet x b . x i can also be expressed as [w i,1 , w i,2 , …, x i, m ] , where w i,j represents the j-th word in tweet x i , and m is the number of words in tweet x i . First, the tweets are divided at the sentence level and then fed into a pre-trained BERT model to obtain their word vector representations, and each sentence corresponds to a word vector matrix. Then, the word-level word vectors are fed into a Bi-LSTM model to learn their context relationships. For word w i,j , in the forward LSTM, the semantic information carried by the previous word w i,j-1 will be passed to it. After learning the semantics of w i,j-1 and the words before it, some unimportant information is forgotten and then passed to the next word w i,j+1 . In the reverse LSTM, the information of the word sequence of w i,j+1 and the words after it that w i,j did not obtain in the forward process will also be passed to it with attenuation through w i,j+1 , so that the Bi-LSTM can retain some global dependency information while paying attention to local sequence information. Its formal representation is: in, and The word w i,j+1 Vector representation of word w in the forward and reverse LSTM hidden layers i,j+1 The final representation h i,j For positive indication With the inverse representation Then, it is input into the Attention layer to model the importance of each word in the sentence, calculate its attention coefficient, and obtain the vector representation of the entire sentence by weighted summation with the content vector. For the hidden vector h of the a-1th layer that has been calculated i,j , the word w in the next layer i,j The attention coefficient of the position e i,j The calculation method is: e i,j =Attention(s a-1 ,h i,j ) After that, the obtained attention coefficient is normalized through the Softmax layer to calculate the word w i,j The attention weight a i,j : Calculate tweet x by weighted sum i The final representation c i : Finally, we get the text content encoding matrix CE=C={c1,c2,…,c n }; (4-2) According to the user forwarding network U = {V′, E′} obtained in step (2), first, the user forwarding network is converted into the corresponding complete Laplacian matrix to obtain the global position code PE Global , the formula is as follows: PE Global =I-D -1 / 2 AD 1 / 2 =Y T ΛY Among them, A is the complete network adjacency matrix, D is the degree matrix of the complete network, Λ and Y correspond to eigenvalues ​​and eigenvectors respectively; Then, since the user retweet network is formed by the fusion of the propagation subgraphs converted from the original propagation tree, the original propagation tree naturally forms a local cluster. Therefore, the local cluster, i.e., the user retweet network, is used to sort the intimacy, i.e.: ON Local =P(v) Among them, P(v) means that for the target node v, other nodes are arranged in descending order according to their intimacy scores calculated with v, and the intimacy distance between the two nodes is obtained, that is, the intimacy matrix, thereby obtaining the local position code PE Local ; Then, the release time attribute of the tweet is used to determine the propagation direction relationship between nodes, and the relative position encoding PE is calculated by Floyd's shortest path algorithm. Relative ,Right now: PE Relative =Floyd(A) Wherein, Floyd(A) is the node shortest distance matrix calculated by executing the Floyd algorithm on the complete network adjacency matrix A; Finally, the global position encoding PE Global , local position encoding PE Local , relative position coding for splicing PE Relative , get the complete position encoding PE: PE=PE Global ⊕PE Local ⊕PE Relative Where ⊕ represents the concatenation of feature vectors; (4-3) Based on the structure code SE obtained in step (3), the text content code CE obtained in step (4-1), and the position code PE obtained in step (4-2), the content code of each tweet node in the text content code is represented as The features of nodes and variables are mapped to d dimensions through linear transformation. The specific calculation formula is: in, is the feature of node i, is the feature of the edge between node i and node j, A 0 and B 0 is the projection matrix, a 0 and b 0 is the bias parameter of the linear mapping; Similarly, the position and structure encodings are linearly mapped to the same dimension and added to the node features: in, is the feature encoding of the position encoding matrix, is the feature encoding of the structure encoding matrix, C 0 and D 0 is the projection matrix, c 0 and d 0 is the bias parameter of the linear mapping, is the final node initial feature representation; The final node initial representation sequence is input into the Graph Transformer network, and the hidden state after the last iteration is taken as the propagation node representation. The Graph Transformer of each layer contains the following parts: self-attention mechanism, feature aggregation, normalization and residual link. For the node update of a certain layer, the specific formula is: Among them, || represents connection, Represents the attention weight, and the specific calculation formula is: After that, the output result is passed to a single-layer feedforward neural network, and the residual connection and normalization are calculated to obtain the embedded representation of the node: For the output of the last layer We refer to it as the propagation node representation Z, namely:

6. The rumor detection method based on information propagation structure according to claim 1 is characterized by: Step (5) specifically includes: According to the propagation node representation Z obtained in step (4), the authenticity of the tweet is scored using a multi-layer perceptron MLP. The specific calculation formula is: y t =softmax(W Y (σ(W z ·Z))) Among them, W Y With W Z is the weight matrix of MLP, Z is the final node representation, σ is the ReLU activation function, for y t , the classification label corresponding to the item with the highest authenticity score is the authenticity label of the predicted information.

Citation Information

Patent Citations

  • A rumor detection method and system based on multi-layer coding network

    CN114328843B

  • Multi-modal rumor detection method and system fusing multi-granularity features

    CN119166907A

  • Rumor detection method based on events and propagation structure

    CN113343126A

  • Game data processing method and device and electronic equipment

    CN116943234A

  • Rumor detection method based on graph attention network

    CN117112786A

Cited By

  • Microblog rumor detection method and system based on static and dynamic knowledge enhancement

    CN121744044A