A social platform rumor forwarding prediction method integrating graph neural network and dual BERT model
By integrating graph neural network and dual Bert model, multiple characteristics of social media users are extracted and aggregated, the problem of difficult to predict rumors spread in the prior art is solved, and accurate prediction of rumors forwarding behavior and identification of potential communication trends are achieved.
Patent Information
- Application Number
- CN202410449730.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-04-15
AI Technical Summary
The existing technology involves prediction of rumor forwarding behaviors of social media users, mainly aiming at the analysis of forwarding behaviors of ordinary tweets, making it difficult to effectively predict the spread trend of rumors.
Using the method of fusion graph neural network and dual Bert model, the GAT layer designs aggregation and splicing of user node features by extracting the text embedding features, user attributes, historical behaviors and topological structure features of user nodes by extracting the text embedding features, user attributes, historical behaviors and topological structure features of user nodes to be forwarded, and outputs the probability of user forwarding rumor tweets.
Accurate predictions of the forwarding of rumors by social network users, can promptly detect potential rumors spread trends, and help governments and relevant institutions take action more quickly to limit the spread of rumors.
Smart Images

Figure CN118247070B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of rumor forwarding prediction on social networks, and more particularly to a method for predicting rumor forwarding on social platforms that integrates a graph neural network and a dual-BERT model. Background Art
[0002] In recent years, researchers have made great efforts in the macro-level rumor propagation problem, such as diffusion model construction, rumor content analysis, and propagation process analysis. However, when social media users receive rumor messages, they will decide whether to spread the rumor due to the influence of rumor content attributes and individual behavioral preferences. Therefore, individual behavior is an important internal motivation that affects the breadth and depth of the rumor propagation cascade. Current research shows that in the social system constructed by social media, individuals have the dual identities of netizens and real social beings. Therefore, individual behavior has psychological and sociological characteristics and is easily affected by the spiral of silence effect. Therefore, some researchers have established multiple models to predict individual forwarding behavior in social media by analyzing the factors that affect individual forwarding behavior in social media.
[0003] Individual forwarding behavior is an important internal motivation that affects the rumor propagation mechanism. Under the influence of information content, social relationships and network environment, individual rumor propagation behavior is obviously different from non-rumor propagation behavior. However, existing prediction models rarely involve the rumor forwarding behavior of social media users, and are mainly focused on the forwarding behavior analysis of ordinary tweets. In view of this, the inventors of this case focus on how to use graph neural networks and Bert models to integrate multiple key factors such as information content, user attributes, topological structure and historical behavior to predict individual users' forwarding behavior of rumors. By predicting rumor forwarding behavior, potential rumor propagation trends can be discovered in a timely manner, allowing government platforms and relevant agencies to take action more quickly to limit the spread of rumors. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a social platform rumor forwarding prediction method that integrates graph neural network and dual BERT model.
[0005] The technical solution adopted by the present invention to solve the technical problem comprises the following steps:
[0006] Step 1: Given a social network G = (V, E), V represents the set of user nodes, and E represents the set of edges between users. Collect each historical forwarded tweet of the user to form a set T = {T1, T2, ..., T N}, N represents the number of users, T i Represents user v i A collection of historical retweets. Represents a rumor to be forwarded, v jRepresents the user who posted the rumor tweet, v i , v j ∈V.
[0007] Step 2: Target the rumor tweets to be forwarded Extract the text embedding features x of the rumor tweets to be forwarded through the Bert+Meanpooling model q . For user v i The historical collection of retweets T i ={t i1 , t i2 , ..., t iM}, and the text embedding feature x is extracted through the Bert+Weightedpooling model i . q , x i and |x q -x i |Concatenate and linearly map to obtain the mapping vector X[i].
[0008] Step 3: Extract user v i The attribute characteristics of H c [i] Historical behavior characteristics H r [i←j], H r [i→j] and topological structure feature H s [i].
[0009] Step 4: Design the GAT layer to perform i Aggregate the attribute characteristics, historical behavior characteristics and topological structure characteristics of neighbor nodes, and concatenate the aggregated information to obtain the concatenated vector Ψ[i].
[0010] Step 5: Concatenate the concatenated vector ψ[i] and the mapped vector X[i] again, input the fully connected layer with Softmax activation function, and output the user v i Retweet a rumor Probability
[0011] Step 6: Set the cross entropy loss function The Adam optimization method is used to train the model. After the learning is completed, the model parameters are used to predict the user to be predicted, and it is determined whether the user will forward the predicted microblog.
[0012] The specific implementation process of step 2 is as follows:
[0013] 2-1. Extract rumor tweets to be forwarded using the Chinese BERT pre-trained model The embedding vector of each word in , and then use the Meanpooling method to output the mean of each word, and finally the rumor tweets to be forwarded The embedding vector is represented as
[0014] 2-2. For user v i The historical collection of retweets T i ={t i1 , t i2 ,…,t iM}, follow step 2-1 to extract the embedding vector {x i1 , x i2 , ..., x iM}.
[0015] 2-3. Extract user v using WeightedPooling method i The historical collection of retweets T i The corresponding embedding vector x i , specifically expressed as:
[0016] x i =WeiAvgPool([μ i1 x i1 , ..., μ iM x iM ]) (1)
[0017] Among them, μ im Represents the weight of each historical retweet:
[0018]
[0019] In the above formula (2), c m Represents user v i and user v j The number of common neighbors of user v j Retweet posted im And by user v i Forward.
[0020] 2-4. x q 、x i and |x q -x i |Concatenate to get the tweet text embedding vector X′[i]. |x q -x i | represents element wise difference.
[0021] 2-5. Add linear mapping to X′[i] to get mapping vector X[i]:
[0022] X[i]=W x X′[i] (3)
[0023] Among them, W x Represents training parameters.
[0024] The specific implementation process of step 3 is as follows:
[0025] 3-1. Extract user v i The attribute features include gender, age, location, and personal description. One-hot encoding is used to represent gender, age, and location information, and the personal description is converted into a vector representation through the Word2Vec model. The above attribute features are concatenated to obtain the user attribute vector C[i], and C[i] is linearly mapped to obtain the attribute feature H through the following formula: c [i]:
[0026] H c [i] = W c C[i] (4)
[0027] Among them, W c Represents training parameters.
[0028] 3-2. Extract user v i and user v j The historical forwarding behavior vector H r [i→j] and H r [i←j]. Define l time intervals and record the number of users v in each time interval. i Forwarding user v j The number of times is normalized and recorded as vector H r [i→j]. Similarly, record the user v in each time interval j Forwarding user v i The number of times is normalized and recorded as vector H r [i←j].
[0029] 3-3. Using SDNE (Structural Deep Network Embedding) to obtain user v i The topological structure characteristics of H s [i] The loss function of SDNE is defined as follows:
[0030] Loss mix =αLoss1+Loss2+yLoss reg (5)
[0031] Among them, Loss1 represents the first-order similarity, Loss2 represents the second-order similarity, and Loss regrepresents the regularization term, α and γ represent the coefficients of the first-order similarity and the regularization term, respectively.
[0032] The specific implementation process of step 4 is as follows:
[0033] 4-1. In the graph attention neural network, for adjacent user pairs (v i , v j ) Calculate the aggregation coefficient on the attribute features
[0034]
[0035]
[0036] In the above formula, N(v i ) represents user v i Neighbor nodes, || represents the splicing operation, Represents the weight parameter The transpose of W c Represents the training parameters. c [j] represents user v j Attribute characteristics.
[0037] 4-2. For adjacent user pairs (v i , v j ) Calculate the aggregation coefficient on historical behavior characteristics
[0038]
[0039]
[0040] In the above formula, N(v i ) represents user v i Neighbor nodes, || represents the splicing operation, Represents the weight parameter The transpose of W r Represents the training parameters.
[0041] 4-3. For adjacent user pairs (v i , v j ) Calculate the aggregation coefficient on the topological structure features
[0042]
[0043]
[0044] In the above formula, N(v i ) represents user v i Neighbor nodes, || represents the splicing operation, Represents the weight parameter The transpose of W s Represents the training parameters.
[0045] 4-4. According to the aggregation coefficient and Calculate user v i The aggregate value of the attribute characteristics, historical behavior characteristics and topological structure characteristics of:
[0046]
[0047]
[0048]
[0049] 4-5. Set user v i Aggregate values of each feature and user v i The original feature values of user v are concatenated to obtain i The final aggregate feature Ψ[i] of is:
[0050] Ψ[i]=T c [i]||T s [i]||T r [i]||H c [i]||H s [i]||H r [i] (15)
[0051] The specific implementation process of step 5 is as follows:
[0052] 5-1. Based on X[i] calculated in step 2 and ψ[i] calculated in step 4, concatenate the two to get Φ[i]:
[0053] Φ[i]=X[i]||ψ[i] (16)
[0054] 5-2. Input Φ[i] into the fully connected layer with Softmax activation function and output user v i Probability of retweeting a rumor
[0055]
[0056] Among them, W FC represents the training vector.
[0057] The specific implementation process of step 6 is as follows:
[0058] 6-1. We define all parameters in this invention as Θ, Γ represents the training set Represents user vi Retweet a rumor data pairs. K represents the number of training sets. Represents the actual forwarding label, and the model is trained using the cross entropy loss function. The rumor forwarding problem of the present invention belongs to a binary classification problem, and the cross entropy loss function calculation formula is as follows:
[0059]
[0060] The present invention utilizes the Adam optimizer to perform minimization learning on the above objective function.
[0061] 6-2. After learning, use the model parameters to predict the user to be predicted, and determine whether the user will forward the rumor tweet to be predicted. In specific implementation, given a new rumor tweet to be predicted And the user v to be predicted i First, extract the content features X[i] of rumor tweets and user tweets, and then obtain the user embedding vector Ψ[i] after aggregation of attribute features, historical forwarding behavior features, and topological structure features based on the trained model. Calculate according to formula (16) and formula (17): The value of The value of user v i Will you retweet rumors?
[0062] The beneficial effects of the present invention are as follows:
[0063] The present invention predicts the user's forwarding behavior by integrating multiple key factors such as rumor tweets, text information of the user's historical forwarded tweets, user topology, attribute information, and historical forwarding behavior. The present invention not only takes into account the impact of content information on forwarding behavior, but also combines the impact of the user's topology, attributes, and historical forwarding behavior on the forwarding intention of the rumor tweet. The present invention designs a dual Bert model to extract information from rumor tweets and user historical tweets. Taking into account that the user's behavior is easily affected by the surrounding users during the rumor forwarding process, the present invention designs a GAT model to aggregate and splice multiple features respectively, which can extract the influence of neighbor users on forwarding behavior in a more fine-grained manner. Given a specific rumor tweet, the present invention can determine whether a specific user of the social network will forward the rumor tweet. In other words, the task of the present invention is to give a rumor to be predicted. and a user v i , predict the corresponding forwarding label
[0064] The spread of rumors on the Internet is often accompanied by malicious attacks, online fraud, privacy leaks and other problems. The present invention not only helps to identify malicious spreads that may threaten network security and privacy at an early stage and take corresponding preventive measures, but also helps to understand the spread mechanism of rumors in Weibo social networks, which has a wide range of applications in public opinion monitoring, enterprise decision-making assistance and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 Diagram of rumor forwarding prediction framework integrating graph neural network and Bert model;
[0066] Figure 2 Diagram of the framework for extracting embedding vectors for tweet text;
[0067] Figure 3 This is a user feature aggregation framework diagram; DETAILED DESCRIPTION
[0068] The present invention will be further described below in conjunction with the accompanying drawings.
[0069] like Figure 1 As shown in the figure, a social platform rumor forwarding prediction method integrating graph neural network and Bert model includes the following steps:
[0070] Step 1: Given a social network G = (V, E), V represents the set of user nodes, and E represents the set of edges between users. Collect each historical forwarded tweet of the user to form a set f = {T1, T2, ..., T N}, N represents the number of users, T i Represents user v i A collection of historical retweets. Represents a rumor to be forwarded, v j Represents the user who posted the rumor tweet, v i , v j ∈V.
[0071] Step 2: Target the rumor tweets to be forwarded Extract the text embedding features x of the rumor tweets to be forwarded through the Bert+Meanpooling model q . For user v i The historical collection of retweets T i ={t i1 , t i2 , ..., t iM}, and the text embedding feature x is extracted through the Bert+Weightedpooling model i . q , x i and |x q -xi |Concatenate and linearly map to obtain the mapping vector X[i].
[0072] Step 3: Extract user v i The attribute characteristics of H c [i] Historical behavior characteristics H r [i←j], H r [i→j] and topological structure feature H s [i].
[0073] Step 4: Design the GAT layer to perform i The attribute characteristics, historical behavior characteristics and topological structure characteristics of its neighboring nodes are aggregated, and the aggregated information is spliced to obtain the splicing vector Ψ[i].
[0074] Step 5: Concatenate the concatenated vector ψ[i] and the mapped vector X[i] again, input the fully connected layer with Softmax activation function, and output the user v i Retweet a rumor Probability
[0075] Step 6: Set the cross entropy loss function The Adam optimization method is used to train the model. After the learning is completed, the model parameters are used to predict the user to be predicted, and it is determined whether the user will forward the predicted microblog.
[0076] The specific implementation process of step 2 is as follows:
[0077] 2-1. Extract rumor tweets to be forwarded using the Chinese BERT pre-trained model The embedding vector of each word in , and then use the Meanpooling method to output the mean of each word, and finally the rumor tweets to be forwarded The embedding vector is represented as
[0078] 2-2. For user v i The historical collection of retweets T i ={t i1 , t i2 , ..., t iM}, follow step 2-1 to extract the embedding vector {x i1 , x i2 , ..., x iM}.
[0079] 2-3. Extract user v using WeightedPooling method i The historical collection of retweets T i The corresponding embedding vector xi , specifically expressed as:
[0080] x i =WeiAvgPool([μ i1 x i1 , ..., μ iM x iM ]) (1)
[0081] Among them, μ im Represents the weight of each historical retweet:
[0082]
[0083] In the above formula (2), c m Represents user v i and user v j The number of common neighbors of user v j Retweet posted im And by user v i Forward.
[0084] 2-4. x q 、x i and |x q -x i | concatenate to get the tweet text embedding vector X′[i], such as Figure 2 As shown; |x q -x i | represents element wise difference.
[0085] 2-5. Add linear mapping to X′[i] to get mapping vector X[i]:
[0086] X[i]=W x X′[i] (3)
[0087] Among them, W x Indicates the training parameters. In specific implementation, the dimension of X[i] is set to 256.
[0088] The specific implementation process of step 3 is as follows:
[0089] 3-1. Extract user v i The attribute features include gender, age, location, and personal description. One-hot encoding is used to represent gender, age, and location information, and the personal description is converted into a vector representation through the Word2Vec model. The above attribute features are concatenated to obtain the user attribute vector C[i], and C[i] is linearly mapped to obtain the attribute feature H through the following formula: c [i]:
[0090] H c [i] = W c C[i] (4)
[0091] Among them, W c In practice, H c The dimension of [i] is set to 64.
[0092] 3-2. Extract user v i and the historical forwarding behavior vector H between user vj r [i→j] and H r [i←j], user v j represents the user who posted the rumor tweet. Define l time intervals and record the number of users v in each time interval. i Forwarding user v j The number of times is normalized and recorded as vector H r [i→j]. Similarly, record the user v in each time interval j Forwarding user v i The number of times is normalized and recorded as vector H r [i←j]. In the specific implementation, the number of time intervals we divide is 64.
[0093] 3-3. Using SDNE (Structural Deep Network Embedding) to obtain user v i The topological structure characteristics of H s [i] The loss function of SDNE is defined as follows:
[0094] Loss mix =αLoss1+Loss2+yLoss reg (5)
[0095] Among them, Loss1 represents the first-order similarity, Loss2 represents the second-order similarity, and Loss reg represents the regularization term, α and γ represent the coefficients of the first-order similarity and the regularization term, respectively.
[0096] For the training of SDNE, we set the number of epochs to 100, the dropout to 0.5, and the number of neurons in the first and second layers of the autoencoder to 128 and 64, respectively. The batch size is configured to be 128. α = 0.01, γ = 0.9.
[0097] The specific implementation process of step 4 is as follows:
[0098] 4-1. In the graph attention neural network, for adjacent user pairs (v i , v j) Calculate the aggregation coefficient on the attribute features
[0099]
[0100]
[0101] In the above formula, N(v i ) represents user v i Neighbor nodes, || represents the splicing operation, Represents the weight parameter The transpose of W c Represents the training parameters. c [j] represents user v j Attribute characteristics.
[0102] 4-2. For adjacent user pairs (v i , v j ) Calculate the aggregation coefficient on historical behavior characteristics
[0103]
[0104]
[0105] In the above formula, N(v i ) represents user v i Neighbor nodes, || represents the splicing operation, Represents the weight parameter The transpose of W r Represents the training parameters.
[0106] 4-3. For adjacent user pairs (v i , v j ) Calculate the aggregation coefficient on the topological structure features
[0107]
[0108]
[0109] In the above formula, N(v i ) represents user v i Neighbor nodes, || represents the splicing operation, Represents the weight parameter The transpose of W s Represents the training parameters.
[0110] 4-4. According to the aggregation coefficient and Calculate user v iThe aggregate value of the attribute characteristics, historical behavior characteristics and topological structure characteristics of:
[0111]
[0112]
[0113]
[0114] 4-5. If Figure 3 As shown, user v i Aggregate values of each feature and user v i The original feature values of user v are concatenated to obtain i The final aggregate feature Ψ[i] of is:
[0115] Ψ[i]=T c [i]||T s [i]||T r [i]||H c [i]||H s [i]||H r [i] (15)
[0116] The specific implementation process of step 5 is as follows:
[0117] 5-1. Based on X[i] calculated in step 2 and Ψ[i] calculated in step 4, concatenate the two to get Φ[i]:
[0118] Φ[i]=X[i]||Ψ[i] (16)
[0119] 5-2. Input Φ[i] into the fully connected layer with Softmax activation function and output user v i Probability of retweeting a rumor
[0120]
[0121] Among them, W FC represents the training vector.
[0122] The specific implementation process of step 6 is as follows:
[0123] 6-1. We define all parameters in this invention as Θ, Γ represents the training set Represents user v i Retweet a rumor data pairs. K represents the number of training sets. Represents the actual forwarding label, and the model is trained using the cross entropy loss function. The rumor forwarding problem of the present invention belongs to a binary classification problem, and the cross entropy loss function calculation formula is as follows:
[0124]
[0125] The present invention uses the Adam optimizer to perform minimization learning on the above objective function. In specific implementation, the learning rate is set to 0.001, the momentum parameters betal and beta2 are set to 0.9 and 0.999 respectively, and epsilon is set to the default value.
[0126] 6-2. After learning, use the model parameters to predict the user to be predicted, and determine whether the user will forward the rumor tweet to be predicted. In specific implementation, given a new rumor tweet to be predicted And the user v to be predicted i First, extract the content features X[i] of rumor tweets and user tweets, and then obtain the user embedding vector Ψ[i] after aggregation of attribute features, historical forwarding behavior features, and topological structure features based on the trained model. Calculate according to formula (16) and formula (17): The value of The value of user v i Will you retweet rumors?
[0127] Experimental data:
[0128] The present invention collects a rumor database mainly based on the Sina platform. The ratio of positive and negative samples is set to 1:1, and the learning rate is set to 0.01, the epoch is set to 100, and the batch size is set to 50. Rumor data information with 5 topics and their propagation relationship network are sorted out from the Sina platform, including users who spread rumors and their corresponding multi-hop neighbors, user attribute information, user tweet history information, etc., to construct a rumor forwarding training database. Using the above information, the rumor tweet information, the historical tweet characteristics of the user to be predicted, the user attribute characteristics, the topological structure characteristics, and the historical forwarding behavior characteristics are analyzed. After 5-fo1d cross-validation, the accuracy rate obtained by the present invention reached 85.24%. Due to the lack of relevant research on rumor forwarding, we use the common tweet forwarding method RLGAT (Representation Learning and GraphAttention Networks) to compare with the present invention. RLGAT uses text content, network structure and social attributes to predict the forwarding of ordinary tweets by users. The specific comparison results are shown in Table 1 below.
[0129] method Accuracy F1-score RLGAT 80.37% 78.64% The present invention 85.24% 83.17%
[0130] From the data in Table 1, it can be seen that the method proposed in the present invention is better than RLGAT, which is 4.53% higher in F1-score and 4.87% higher in Accuracy.
Claims
1. A social platform rumor forwarding prediction method integrating graph neural network and dual BERT model, characterized by The steps include: Step 1: Given a social network G = (V, E), where V represents the set of user nodes and E represents the set of edges between users; collect each historical forwarded tweet of the user to form a set T = {T1, T2, ..., T N }, N represents the number of users, T i Represents user v i A collection of historical retweets; Represents the rumor tweet to be forwarded, v j Represents the user who posted the rumor tweet, v i , v j ∈V; Step 2: Target the rumor tweets to be forwarded The text embedding feature x of the rumor tweet to be forwarded is extracted through the Bert+Meanpooling model q ; For user v i The historical collection of retweets T i ={t i1 , t i2 , ..., t iM }, and the text embedding feature x is extracted through the Bert+Weightedpooling model i ; x q , x i and |x q -x i |Concatenate and linearly map to obtain the mapping vector X[i]; Step 3: Extract user v i The attribute characteristics of H c [i] Historical behavior characteristics H r [i] and topological structure characteristics H s [i]; Step 4: Design the GAT layer to perform i The attribute characteristics of H c [i] Historical behavior characteristics H r [i] and topological structure characteristics H s [i] is aggregated, and the aggregated information is concatenated to obtain the concatenated vector Ψ[i]; Step 5: Concatenate the concatenated vector ψ[i] and the mapped vector X[i] again, input the fully connected layer with Softmax activation function, and output the user v i Retweet a rumor Probability Step 6: Set the cross entropy loss function The model is trained using the Adam optimization method. After learning, the model parameters are used to predict the user to be predicted, and it is determined whether the user will forward the predicted microblog. The specific implementation process of step 3 is as follows: 3-1. Extract user v i The attribute features include gender, age, geographic location and personal description; use one-hot encoding to represent gender, age and geographic location information, and convert personal description into vector representation through the Word2Vec model; concatenate the above attribute features to obtain the user attribute vector C[i], and linearly map C[i] to obtain the attribute feature H through the following formula c [i]: H c [i]=W c C[i] (4) Among them, W c represents the training parameters; 3-2. Extract user v i and user v j The historical forwarding behavior vector H r [i→j] and H r [i←j], user v j represents the user who posted the rumor tweet; define l time intervals and record the number of users v in each time interval i Forwarding user v j The number of forwardings is normalized and recorded as vector H r [i→j]; also record the user v in each time interval j Forwarding user v i The number of forwardings is normalized and recorded as vector H r [i←j]; H r [i→j] and H r [i←j] is concatenated to obtain the historical behavior feature H r [i]; 3-3. Using SDNE to obtain user v i The topological structure characteristics of H s [i]; The loss function of SDNE is defined as follows: Loss mix =αLoss1+Loss2+γLoss reg (5) Among them, Loss1 represents the first-order similarity, Loss2 represents the second-order similarity, and Loss reg represents the regularization term, α and γ represent the coefficients of the first-order similarity and the regularization term, respectively; The specific implementation process of step 4 includes: 4-1. In the graph attention neural network, for adjacent user pairs (v i , v j ) Calculate the aggregation coefficient on the attribute features In the above formula, N(v i ) represents user v i Neighbor nodes, || represents the splicing operation, Represents the weight parameter The transpose of W c represents the training parameters; H c [j] indicates the attribute features used for j; The specific implementation process of step 4 includes: 4-2. For adjacent user pairs (v i , v j ) Calculate the aggregation coefficient on historical behavior characteristics In the above formula, N(v i ) represents user v i Neighbor nodes, || represents the splicing operation, Represents the weight parameter The transpose of W r represents the training parameters; The specific implementation process of step 4 includes: 4-3. For adjacent user pairs (v i , v j ) Calculate the aggregation coefficient on the topological structure features In the above formula, N(v i ) represents user v i Neighbor nodes, || represents the splicing operation, Represents the weight parameter The transpose of W s represents the training parameters; The specific implementation process of step 4 includes: 4-4. According to the aggregation coefficient and Calculate the aggregate values of attribute features, historical behavior features, and topological structure features: The specific implementation process of step 4 includes: 4-5. Concatenate the aggregated values of each feature and the original feature value of the user to obtain the user v i The final aggregate feature Ψ[i] of is: Ψ[i]=T c [i]||T s [i]||T r [i]||H c [i]||H s [i]||H r [i] (15)。 2. The social platform rumor forwarding prediction method integrating graph neural network and dual BERT model according to claim 1 is characterized in that The specific implementation process of step 2 is as follows: 2-1. Extract rumor tweets to be forwarded using the Chinese BERT pre-trained model The embedding vector of each word in , and then use the Meanpooling method to output the mean of each word, and finally the rumor tweets to be forwarded The embedding vector is represented as 2-2. For user v i The historical collection of retweets T i ={t i1 , t i2 ,…,t iM }, follow step 2-1 to extract the embedding vector {x i1 , x i2 , ..., x iM }; 2-3. Extract user v using WeightedPooling method i The historical collection of retweets T i The corresponding embedding vector x i , specifically expressed as: x i =WeiAvgPool([μ i1 x i1 ,...,μ iM x iM ]) (1) Among them, μ im Represents the weight of each historical retweet: In the above formula (2), c m Represents user v i and user v j The number of common neighbors of user v j Retweet posted im And by user v i Forward; 2-4. x q 、x i and |x q -x i |Concatenate to get the tweet text embedding vector X′[i]; |x q -x i | represents the corresponding difference of elements; 2-5. Add linear mapping to X′[i] to get mapping vector X[i]: X[i]=W x X′[i] (3) Among them, W x Represents training parameters.
3. The social platform rumor forwarding prediction method integrating graph neural network and dual BERT model according to claim 2 is characterized in that The specific implementation process of step 5 is as follows: 5-1. Based on X[i] calculated in step 2 and Ψ[i] calculated in step 4, concatenate the two to get Φ[i]: Φ[i]=X[i]||ψ[i] (16) 5-2. Input Φ[i] into the fully connected layer with Softmax activation function to output user v i Probability of retweeting a rumor Among them, W FC represents the training vector.
4. The social platform rumor forwarding prediction method integrating graph neural network and dual BERT model according to claim 3 is characterized in that The specific implementation process of step 6 is as follows: 6-1. Define all parameters as Θ, Γ represents the training set Represents user v i Retweet a rumor data pairs, K represents the number of training sets, Represents the actual forwarding label. The model is trained using the cross entropy loss function. The cross entropy loss function calculation formula is as follows: 6-2. After learning, use the model parameters to predict the user to be predicted, and determine whether the user will forward the rumor tweet to be predicted.
Citation Information
Patent Citations
Social media rumor detection method based on network information propagation graph modeling
CN112199608A
Rumor detection method and device based on graph neural network feature aggregation
CN113139052A