Rumor detection method based on dual-domain perception structure feature fusion learning
By constructing social media network graphs and using graph attention networks to extract features, combining the propagation density peak and slope division method to identify key time windows, and using the mutual attention mechanism to fusion features, the problems of insufficient mining of communication feature and insufficient fusion of multidimensional features in the existing rumor detection methods are solved, and accurate rumor detection and improved detection performance are achieved.
Patent Information
- Application Number
- CN202510756596.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing rumor detection methods are not fully mined for transmission characteristics in key time periods, it is difficult to effectively model the global structure of information dissemination, and the multi-dimensional feature fusion is insufficient, and the joint modeling ability of user-information interaction dynamics is lacking, which affects detection performance.
By constructing a post dissemination network diagram and a user social network diagram, using the graph attention network to extract features, combining the propagation density peak and slope division method to identify key time windows, using a mutual attention mechanism to integrate post dissemination characteristics and user social characteristics, constructing a projection matrix for weighted summing, performing homogeneous interactive information modeling, and finally inputting the rumor detection module for classification.
Accurate rumor detection is realized, which improves detection accuracy and robustness, especially after integrating social interaction characteristics, it significantly improves detection performance, and is suitable for real-time detection of large-scale social media data.
Smart Images

Figure CN120277541A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of natural language processing and information dissemination, and particularly relates to a rumor detection method based on dual-domain perception structure feature fusion learning. Background Art
[0002] With the rapid development of the Internet and social media, the public can obtain information and exchange views more efficiently. Different from traditional rumor dissemination channels, rumors on social media break through geographical and time boundaries and spread rapidly. However, rumors in social media have unique dissemination patterns. How to effectively detect them using these characteristics is crucial for purifying the network environment and maintaining information security. The rapid development of deep learning technology provides important support for rumor detection. Researchers extract features from multiple dimensions such as text content, user characteristics, and dissemination paths to construct high-dimensional event representations and evaluate information authenticity. For example, recurrent neural networks (RNN, LSTM, GRU) can capture the time dependence of information dissemination due to their excellent time series modeling capabilities; convolutional neural networks (CNN) are widely used in text analysis and dissemination pattern modeling due to their strong feature extraction capabilities. However, CNN mainly focuses on local features and is difficult to effectively model the global structure of information dissemination. GNN can effectively capture the complex dynamic features of information dissemination by modeling nodes and their relationships in the graph structure, showing good application prospects in rumor detection.
[0003] To accurately detect rumors, researchers have deeply analyzed information such as content features (such as text and images) involved in rumors, user attributes, and dissemination structures, and explored methods such as machine learning, tree structure modeling, and graph structure learning. Through sorting out existing research, it is found that existing rumor detection methods still face several key problems: First, the dissemination characteristics in key time periods are not fully mined. Second, rumor dissemination is often complex and sudden. When existing methods deal with the dynamic features in the intensive dissemination stage, they fail to fully utilize limited time series to mine key features, making it difficult to accurately capture the core information of rumor dissemination. Third, the multi-dimensional feature fusion is insufficient. Existing methods split the coupling relationship between user social behavior features and information dissemination structures, lacking the ability to jointly model the user-information interaction dynamics, which affects the detection performance. Summary of the Invention
[0004] To solve the above problems, the present invention provides a rumor detection method based on dual-domain perception structure feature fusion learning.
[0005] To achieve the above object, the present invention is realized through the following technical solutions: The present invention provides a rumor detection method based on dual-domain perception structure feature fusion learning, including the following steps: S1. Obtain a social media message event dataset, construct a post propagation network graph and a user social network graph based on the dataset, and extract post propagation features and user social features respectively through a graph attention network encoder; S2. Use the propagation density peak and slope division method to identify the key time window of rumor diffusion, and extract relevant sub-structure features from the post propagation network graph and the user social network graph; S3. Process the post propagation sub-features and user social sub-features through a mutual attention mechanism to obtain post propagation fusion features and user social fusion features; S4. Construct a projection matrix through the post propagation fusion features and user social fusion features, and then use the projection matrix to perform weighted summation on the post propagation fusion features and user social fusion features to obtain post propagation interaction features and user social interaction features, and further generate the first weighted post propagation features and the first weighted user social features; S5. Perform homogeneous interaction information modeling on the first weighted post propagation features and the first weighted user social features to obtain the final homogeneous interaction comprehensive representation; S6. After splicing the final homogeneous interaction comprehensive representation with the original tweet features, input them into the rumor detection module for classification to obtain the rumor detection result.
[0006] Furthermore, step S1 specifically includes: The social media message event dataset , where represents the th event, is the total number of events; For each event construct a post propagation network graph , the root node of the propagation graph is the event statement ; is the node set of the post propagation network graph, and a node represents a tweet; is the edge set of the post propagation network graph, representing the interaction between tweets; the adjacency matrix of the post propagation network graph is , when , it means that the th tweet and the th tweet have an interaction relationship; when , it means that the th tweet and the th tweet have no interaction relationship; For each event construct a user social network graph , is the node set of the user social network graph, and a node represents a user; It is the edge set of the user social network graph, representing the social relationships between users; the adjacency matrix of the user social network graph is When it means that the th user has a follow relationship with the th user; when it means that the th user has no follow relationship with the th user; Input the post propagation network graph and the user social network graph into the graph attention network GAT encoder for feature extraction to obtain the post propagation feature and the user social feature ; the graph attention network GAT encoder adopts a two-layer graph attention network GAT encoder. The first-layer graph attention network GAT encoder preliminarily updates the node features, and the second-layer graph attention network GAT encoder processes the output features of the first-layer graph attention network GAT encoder.
[0007] Furthermore, step S2 specifically includes: Calculate the propagation density for each event , and the formula is as follows: , where represents the indicator function, which takes the value of 1 when the time information belongs to the time , and 0 otherwise; the time is divided by hours; represents that the node belongs to the node set ; According to the propagation time and propagation density of the event , obtain the propagation density graph, and use the sliding window method to detect the local maximum value to identify the outbreak node. The formula is as follows: , where represents the time point where the local maximum value is located, represents the peak value of the fluctuation during the propagation process, represents the window size, represents the maximum value taking operation; select the two highest peak values of the propagation density, denoted as the first peak time point , and the second peak time point ; Calculate the slope within the neighborhood of each peak in the propagation density graph, and the formula is as follows: , where Represents the density value of the propagation density map at the time point , represents the window length; represents the slope, representing the change trend of the propagation density at this moment; Starting from the local maximum point in the propagation density map corresponding to the peak value of the fluctuation in the propagation process , search for points with a slope of ± to both sides as the boundaries of the outbreak interval. Starting from the first peak time point , scan left and right respectively to find the first time point that satisfies ; Determine the boundaries of the propagation outbreak interval: start time: ; end time: ; Finally, obtain the propagation outbreak interval corresponding to the first peak time point : ; Similarly, obtain the propagation outbreak interval corresponding to the second peak time point ; According to the obtained propagation outbreak interval, filter out the nodes and edges that meet the outbreak interval from the post propagation network diagram to obtain the post propagation sub-diagram , and the post propagation sub-diagram satisfies the following formula: , , wherein, represents the node set of the post propagation sub-diagram, represents the edge set of the post propagation sub-diagram; Through the above process, extract the post propagation sub-diagram of the first peak time point from the post propagation network diagram and the post propagation sub-diagram of the second peak time point ; The post propagation sub-diagram and are respectively subjected to feature extraction through the graph attention network GAT encoder to obtain the corresponding first post propagation sub-feature and the second post propagation sub-feature ; According to the association between the post propagation network diagram and the user social network diagram, extract the user social sub-diagram corresponding to each post propagation sub-diagram from the user social network diagram, and use the graph attention network GAT encoder to obtain the first user social sub-feature and the second user social sub-feature .
[0008] Furthermore, step S3 specifically includes: The first post propagation sub-feature , Second post propagation sub-feature , First user social sub-feature and second user social sub-feature Perform feature mapping through a linear transformation function to obtain the first post propagation mapping sub-feature , Second post propagation mapping sub-feature , First user social mapping sub-feature and second user social mapping sub-feature ; Adopt a multi-head attention mechanism to fuse feature representations. In the multi-head attention mechanism, each attention head learns different attention distributions through independent parameters; for each attention head, calculate the attention weight matrix between the first post propagation mapping sub-feature and the second post propagation mapping sub-feature , and the formula is as follows: , where, represents the parameter of the th attention head, Softmax represents the normalization function, represents the scaling factor; similarly, calculate the attention weight matrix between the first user social mapping sub-feature and the second user social mapping sub-feature ; Perform weighted summation on the first post propagation mapping sub-feature and the second post propagation mapping sub-feature to obtain the post propagation fusion sub-feature, and the formula is as follows: , where, represents the post propagation fusion sub-feature output by the th attention head; similarly, obtain the user social fusion sub-feature ; in the multi-head attention mechanism, multiple attention heads calculate different attention distributions in parallel, splice the outputs of multiple attention heads together, and map them to the final feature space through a linear transformation, and the formula is as follows , , where, represents the post propagation splicing sub-feature, represents the splicing operation, represents the first linear transformation matrix, represents the user social splicing sub-feature, represents the second linear transformation matrix; Propagate and splice the post sub - feature and the post propagation feature are fused to obtain the post propagation fusion feature, which is expressed by the formula as follows: , wherein, represents the post propagation fusion feature; similarly, the user social fusion feature is obtained.
[0009] Furthermore, step S4 specifically includes: Construct a projection matrix through the post propagation fusion feature and the user social fusion feature , and then use the projection matrix to perform weighted summation on the post propagation fusion feature and the user social fusion feature to obtain the post propagation interaction feature and the user social interaction feature , which is expressed by the formula as follows: , , , wherein, represents the hyperbolic tangent activation function, represents the parameter of the projection matrix, represents the linear transformation matrix of the user social feature, represents the linear transformation matrix of the post propagation feature; Perform attention weight assignment on the post propagation interaction feature and the user social interaction feature respectively, calculate the importance of each feature dimension through the Softmax function, and perform weighted summation to obtain the first weighted post propagation feature and the first weighted user social feature : , , wherein, represents the Softmax function.
[0010] Furthermore, step S5 specifically includes: Through the self - attention module, the first weighted post propagation feature and the first weighted user social feature Perform weight redistribution, learn the correlation and dynamic importance of internal features, and obtain the second weighted post propagation feature and the second weighted user social feature , which is expressed by the following formula: , , where, represents the self-attention mechanism; Use the gated recurrent unit GRU to perform temporal modeling on the features and to obtain the final hidden state as , which is expressed by the following formula: , where, represents the gated recurrent unit operation; Concatenate the first weighted post propagation feature , the first weighted user social feature and the final hidden state to obtain the final homogeneous interaction comprehensive representation , which is expressed as: .
[0011] Furthermore, step S6 specifically includes: Concatenate the final homogeneous interaction comprehensive representation with the original tweet feature of the event statement to obtain the final comprehensive feature , which is expressed by the following formula: , Calculate the predicted probability vector through the fully connected layer and the Softmax function, which is expressed by the following formula: , where, represents the fully connected layer operation.
[0012] Furthermore, during the training process, optimize by minimizing the cross-entropy loss between the predicted probability and the true label distribution: , where, represents the cross-entropy loss, is the number of samples in the dataset, represents the number of classes, that is, the total number of classes in the classification task, represents that this is the regularization term, representing the sum of squares of the model parameters , is the regularization factor, represents the event true label distribution.
[0013] The advantages of the present invention are as follows: The method of the present invention constructs a post propagation network and a user social network, extracts cross-domain features using a graph attention network (GAT), and identifies the evolution features of key time windows based on the slope division method to extract high-influence group features. Through the mutual attention mechanism, the post propagation features and user social features are fused, and the heterogeneous and homogeneous interaction modules are combined to capture global dependencies and temporal patterns, realizing accurate rumor detection. Experimental results show that the method of the present invention is superior to existing models on multiple real social media datasets. Especially after fusing social interaction features, the detection accuracy and robustness are significantly improved. In addition, the method of the present invention performs excellently in terms of time complexity and early detection, and is suitable for real-time detection of large-scale social media data. Ablation experiments and hyperparameter sensitivity analysis further verify the effectiveness of the framework, providing optimization guidance for practical applications. Brief Description of the Drawings
[0014] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.
[0015] Figure 1 is the step flowchart of the method of the present invention; Figure 2 is the result of rumor statistical analysis; Figure 2 (a) is the rumor statistical analysis graph on the Twitter15 dataset; Figure 2 (b) is the rumor statistical analysis graph on the Weibo dataset; Figure 3 is the comparison of early rumor detection results on the Twitter15 dataset; Figure 4 is the comparison of early rumor detection results on the Twitter16 dataset; Figure 5 is the comparison of early rumor detection results on the Pheme dataset; Figure 6 is the result of the ablation experiment of the method of the present invention; Figure 7 is the influence comparison of the setting of the number of outbreak section divisions ( ) on the method of the present invention; Figure 8 is the influence comparison of the setting of the proportion of deleted edges ( ) on the method of the present invention; Figure 9 is the setting of the slope threshold ( ), and the impact comparison on the method of the present invention is shown. Specific embodiments
[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] Embodiment 1 In this embodiment, as Figure 1 shown, the present invention provides a rumor detection method based on dual-domain perception structure feature fusion learning. The specific steps include: S1. Obtain a social media message event dataset, construct a post propagation network graph and a user social network graph based on the dataset, and extract post propagation features and user social features respectively through a graph attention network encoder.
[0018] Specifically, the social media message event dataset , where represents the th event, is the total number of events; For each event construct a post propagation network graph , and the root node of the propagation graph is the event statement ; is the node set of the post propagation network graph, and a node represents a tweet; is the edge set of the post propagation network graph, representing the interaction between tweets; the adjacency matrix of the post propagation network graph is , when , it means that the th tweet has an interaction relationship with the th tweet; when , it means that the th tweet has no interaction relationship with the th tweet; For each event construct a user social network graph , is the node set of the user social network graph, and a node represents a user; is the edge set of the user social network graph, representing the social relationship between users; the adjacency matrix of the user social network graph is , when , it means that the A user has a following relationship with the th user; when it is, it means that the th user has no following relationship with the th user; Input the post propagation network graph and the user social network graph into the Graph Attention Network (GAT) encoder for feature extraction to obtain the post propagation features and user social features ; The GAT encoder adopts a two-layer GAT encoder. The first-layer GAT encoder preliminarily updates the node features, and the second-layer GAT encoder processes the output features of the first-layer GAT encoder.
[0019] S2. Use the propagation density peak and slope division method to identify the key time window of rumor diffusion, and extract relevant sub-structure features from the post propagation network graph and the user social network graph.
[0020] Specifically, regard rumor propagation as a dynamic process with time series characteristics. In order to identify the outbreak interval in the propagation process, it is first necessary to divide the propagation behavior of the rumor event by time and calculate its propagation density in different time periods; for the event , the propagation density represents the number of propagation behaviors (such as comments, forwards) participated by users from the propagation node set in this time period. According to the timestamps attached to each propagation record in the original dataset, divide the entire propagation process; Calculate the propagation density for each event , and the formula is as follows: , where represents the indicator function, which takes the value of 1 when the time information belongs to the time , and 0 otherwise; the time is divided by hours; represents that the node belongs to the node set ; According to the propagation time and propagation density of the event , obtain the propagation density graph, and use the sliding window method to detect local maxima and identify outbreak nodes. The formula is as follows: , where represents the time point where the local maximum is located, represents the peak value of the fluctuation during the propagation process represents the window size represents the maximum operation; to ensure the analysis of local propagation outbreaks, this study selects the two peaks with the highest propagation density, denoted as the first peak time point and the second peak time point ; Calculate the slope within the neighborhood of each peak in the propagation density map, and the formula is as follows: ; where represents the density value of the propagation density map at time point ; represents the window length; represents the slope, representing the change trend of the propagation density at this moment; this slope is used to assist in finding the boundaries of the intervals where the density changes significantly.
[0021] Starting from the local maximum point in the propagation density map corresponding to the peak value of the fluctuation during the propagation process , search for points with slopes of ± on both sides as the boundaries of the outbreak interval. Starting from the first peak time point , scan left and right respectively to find the first time point that satisfies ; determine the boundaries of the propagation outbreak interval: start time: ; end time: ; finally, obtain the propagation outbreak interval corresponding to the first peak time point : ; This time period is considered the critical period during which the propagation intensity rises or falls significantly during the event propagation process, providing a clear time boundary basis for the subsequent extraction of propagation substructure characteristics. Similarly, obtain the propagation outbreak interval corresponding to the second peak time point ; According to the obtained propagation outbreak interval, filter out the nodes and edges that meet the outbreak interval from the post propagation network graph to obtain the post propagation subgraph , and the post propagation subgraph satisfies the following formula: ; ; where represents the node set of the post propagation subgraph represents the edge set of the post propagation subgraph; the propagation subgraph selected through this time window is the substructure of the nodes and edges with the highest propagation density and information volume during this time period, and is the basis for the subsequent extraction of structural characteristics.
[0022] Through the above process, the first peak time point is extracted from the post propagation network diagram of the post propagation sub - diagram and the second peak time point of the post propagation sub - diagram ; The post propagation sub - diagrams and are respectively subjected to feature extraction through the graph attention network GAT encoder to obtain the corresponding first post propagation sub - feature and the second post propagation sub - feature ; According to the association between the post propagation network diagram and the user social network diagram, the user social sub - diagrams corresponding to each post propagation sub - diagram are extracted from the user social network diagram, and the graph attention network GAT encoder is used to obtain the first user social sub - feature and the second user social sub - feature .
[0023] S3. The post propagation sub - features and the user social sub - features are processed through the mutual attention mechanism to obtain the post propagation fusion feature and the user social fusion feature.
[0024] Specifically, the first post propagation sub - feature , the second post propagation sub - feature , the first user social sub - feature and the second user social sub - feature are subjected to feature mapping through the linear transformation function to obtain the first post propagation mapped sub - feature , the second post propagation mapped sub - feature , the first user social mapped sub - feature and the second user social mapped sub - feature ; So that the feature vectors from different sources are mapped into the feature space of the same dimension, which is convenient for subsequent mutual attention mechanism learning. The formula is as follows: , , , , Among them, , respectively represent the learned weight matrices for mapping the post propagation sub - features and the user social sub - features, , respectively represent the bias terms of the post propagation sub - features and the user social sub - features; represents the matrix multiplication operation; Ensure that all the mapped features , , , The dimensions are consistent and mapped into a unified feature space. This unified feature space provides a feature alignment basis for the mutual attention mechanism between subsequent heterogeneous structures, ensuring that structural features from different sources can be associated and modeled and fusion-learned in the same semantic space.
[0025] Use the multi-head attention mechanism to fuse feature representations. In the multi-head attention mechanism, each attention head learns different attention distributions through independent parameters; for each attention head, calculate the first post-propagation mapping sub-feature and the second post-propagation mapping sub-feature between the attention weight matrix , the formula is as follows: , where, represents the parameter of the th attention head, Softmax represents the normalization function, represents the scaling factor; similarly, calculate the first user social mapping sub-feature and the second user social mapping sub-feature between the attention weight matrix ; Perform weighted summation on the first post-propagation mapping sub-feature and the second post-propagation mapping sub-feature to obtain the post-propagation fusion sub-feature, the formula is as follows: , where, represents the post-propagation fusion sub-feature output by the th attention head; similarly, obtain the user social fusion sub-feature ; in the multi-head attention mechanism, multiple attention heads calculate different attention distributions in parallel, splice the outputs of multiple attention heads together, and map them to the final feature space through a linear transformation, the formula is as follows , , where, represents the post-propagation splicing sub-feature, represents the splicing operation, represents the first linear transformation matrix, represents the user social splicing sub-feature, represents the second linear transformation matrix; in the above process, and are matrices used for linearly transforming the concatenated post propagation sub - features and user social sub - features respectively. Specifically, these linear transformation matrices are optimized through the backpropagation algorithm during the training process. In the multi - head attention mechanism, after the output of the attention heads is concatenated, the dimension may increase, so it is necessary to use the linear transformation matrices and to map the concatenated features to the target feature space and ensure the consistency of the feature dimensions. These matrices are usually initialized with random values and are gradually adjusted as the model is updated during the training process to learn the best feature representation. The purpose of the linear transformation is to transform the concatenated feature vectors to a feature space more suitable for the subsequent processing of the model through mapping, further enhancing the expressive ability of the model.
[0026] Fuse the post propagation concatenated sub - features and the post propagation features to obtain the post propagation fused features, which is expressed by the formula as follows: , where represents the post propagation fused features; similarly, the user social fused features are obtained.
[0027] S4. Construct a projection matrix through the post propagation fused features and the user social fused features, and then use the projection matrix to perform weighted summation on the post propagation fused features and the user social fused features to obtain the post propagation interaction features and the user social interaction features, and further generate the first weighted post propagation features and the first weighted user social features.
[0028] Specifically, construct a projection matrix through the post propagation fused features and the user social fused features , and then use the projection matrix to perform weighted summation on the post propagation fused features and the user social fused features to obtain the post propagation interaction features and the user social interaction features , which is expressed by the formula as follows: , , , where represents the hyperbolic tangent activation function, which is usually used to map the feature values after linear transformation to a certain range to enhance the non - linear expressive ability. represents the parameter of the projection matrix, which is used to connect the post propagation features and user social characteristics to generate a projection matrix . This is a parameter matrix obtained through learning. A linear transformation matrix representing user social characteristics, used for the user social fusion characteristics for transformation A linear transformation matrix representing post propagation characteristics, used for the post propagation fusion characteristics for transformation; Respectively perform attention weight allocation on the post propagation interaction characteristics and user social interaction characteristics , calculate the importance of each feature dimension through the Softmax function, and perform weighted summation to obtain the first weighted post propagation feature and the first weighted user social feature : , , wherein, represents the Softmax function. It should be noted that the Softmax here is used to allocate cross-dimensional attention weights. Although it is similar in form to the self-attention mechanism, it essentially belongs to heterogeneous interaction modeling between structures and is different from the subsequent homogeneous attention mechanism.
[0029] S5. Perform homogeneous interaction information modeling on the first weighted post propagation feature and the first weighted user social feature to obtain the final homogeneous interaction comprehensive representation.
[0030] Specifically, in order to further explore the dynamic evolution characteristics within the post propagation structure and the user social structure respectively, the present invention performs homogeneity modeling on the weighted features obtained in step S4. First, the self-attention mechanism is introduced to capture the importance distribution within the features, so as to obtain more representative second weighted features; through the self-attention module, the first weighted post propagation feature and the first weighted user social feature are reallocated weights to learn the correlation relationship and dynamic importance of internal features, and the second weighted post propagation feature and the second weighted user social feature are obtained. The formula is as follows: , , wherein, represents the self-attention mechanism; Use the gated recurrent unit GRU for the feature and perform temporal modeling to obtain the final hidden state as , and the formula is as follows: , wherein, represents the gated recurrent unit operation; concatenate the first weighted post propagation feature , the first weighted user social feature and the final hidden state to obtain the final homogeneous interaction comprehensive representation , and the formula is expressed as: ; wherein, is the final representation that comprehensively fuses the dynamic evolution feature and the original structure feature, and serves as the input for subsequent heterogeneous interaction fusion.
[0031] S6. After concatenating the final homogeneous interaction comprehensive representation with the original tweet feature, input it into the rumor detection module for classification to obtain the rumor detection result.
[0032] Specifically, concatenate the final homogeneous interaction comprehensive representation with the event statement and the original tweet feature to obtain the final comprehensive feature , and the formula is as follows: , Calculate the predicted probability vector through the fully connected layer and the Softmax function, and the formula is as follows: , wherein, represents the fully connected layer operation.
[0033] Specifically, during the training process, optimize by minimizing the cross-entropy loss between the predicted probability and the true label distribution: , wherein, represents the cross-entropy loss, is the number of samples in the dataset, represents the number of classes, that is, the total number of classes in the classification task, represents that this is the regularization term, representing the sum of squares of the model parameters , is the regularization factor, represents the event 's true label distribution.
[0034] Example 2 This embodiment conducts an experimental analysis of the method of the present invention: 1. Dataset Introduction and Statistical Analysis (1) Dataset Introduction In the experiment, three publicly available social media datasets are used to evaluate the performance of the method of the present invention. The datasets include the Twitter15 dataset, the Twitter16 dataset, and the Pheme dataset. The statistical analysis of the datasets is shown in Table 1; Table 1 Statistical Analysis of Datasets (2) Statistical Analysis The present invention is based on the volatility characteristics in the process of rumor propagation, extracts the outbreak section, and thus mines the dense feature interval in information dissemination. To verify the feasibility of the method, this article conducts a statistical analysis on two representative datasets: Twitter15 as a typical English social media dataset, and Weibo representing the Chinese social platform. The two are complementary in language structure and user behavior, and have good experimental representativeness and promotion value. The analysis results are as Figure 2 shown, Figure 2 (a) is the statistical analysis chart of rumors on the Twitter15 dataset, Figure 2 (b) is the statistical analysis chart of rumors on the Weibo dataset; the horizontal axis represents the time of rumor propagation, and the vertical axis represents the number of comments or forwards, that is, the propagation density. It can be observed from the figure that the curve of rumor events changing with time shows obvious volatility, presenting characteristic peak fluctuations. These fluctuations not only show rises and falls in time, but also show certain regular rises and falls in different time periods. This law has been fully verified in the statistical results of both the Twitter15 and Weibo datasets. Specifically, the propagation density rises sharply in some time periods, forming a short peak, and then gradually decreases over time. This phenomenon reflects the active interaction of users in the process of information dissemination, as well as the volatility of the information dissemination intensity.
[0035] Through the analysis of the propagation density fluctuations, the outbreak events of information dissemination can be identified. The peak of the propagation density corresponds to the peak of user interaction, and then rapidly decays, reflecting the suddenness of the propagation. This fluctuation characteristic indicates that the information diffusion is not linear growth, but is affected by user attention and dissemination mechanisms, showing phased outbreaks. Dividing the outbreak density interval helps to locate the key nodes of rumor diffusion, extract accurate time features, and further reveal the dynamic pattern of information dissemination. Combining user interaction behavior and propagation path analysis can enhance the perception ability of critical moments, improve the detection accuracy and adaptability, especially in a dynamic social environment, and improve the response speed and early warning ability to outbreak events.
[0036] 2. Evaluation Indicators The present invention uses accuracy and F1 score as the core evaluation criteria to comprehensively measure the overall performance of the method of the present invention in the rumor detection task. The advantage of the F1 score is that it comprehensively considers the precision and recall ability of classification, ensuring that the model has high accuracy when predicting positive class samples and can cover all positive class instances as much as possible. In the rumor detection task, it can effectively make up for the possible biases when only using accuracy for evaluation, enabling the model to still maintain good discriminative ability when facing the problem of class imbalance. In the experiment, comprehensively analyzing the accuracy and F1 score helps to comprehensively measure the model performance and provide a reference basis for algorithm optimization.
[0037] 3. Experimental settings The default optimization configurations of each comparative method are adopted, implemented under the PyTorch framework, and Adam optimization is selected. The present invention uses the processed user social and post interaction dataset and constructs a propagation relationship graph. The specific experimental parameters are as follows: the learning rate is set to 0.001, the batch size is 128, and the dropout rate is 0.2. The text nodes are initialized with TF-IDF (dimension 5000), and the user features use standard initial vectors. The output dimensions of both the graph attention network (GAT) and the fusion gating unit are 64, and the number of heads of the multi-head attention is 4. Outbreak section Select to include 2 sub-structures, slope thresholds Set to 0.4, 0.4, and 0.5; the edge deletion ratio Are 0.2, 0.3, and 0.3 respectively. The edge deletion ratio q is a regularization and data augmentation strategy in the graph representation learning process, used to perturb the graph structure during the training stage to improve the generalization ability of the model, prevent overfitting, and enhance the robustness to noisy edges. This operation constructs diverse graph structure views by randomly deleting some edges at the ratio q, which helps the model learn more stable and discriminative representations. The early stopping strategy is adopted during the training process. When the validation set loss does not decrease for 10 consecutive rounds, the training stops, and five-fold cross-validation is performed to improve the experimental robustness. Finally, the performance of the test set is evaluated using the best parameters obtained from the validation set.
[0038] 4. Comparative methods The method of the present invention is compared with several advanced comparative models. The comparative models include: DTC model: A detection method based on a decision tree classifier, which classifies and discriminates by relying on artificially constructed text features.
[0039] RFC model: A detection method based on a random forest classifier, which detects by fusing user temporal behavior and event association features.
[0040] SVM-TK Model: A detection method based on linear support vector machine (SVM), which improves the discriminant performance by constructing propagation time series features.
[0041] RvNN Model: A detection method based on recurrent neural network (RNN), which optimizes the classification effect by modeling the temporal dependence in the post propagation process.
[0042] PLAN Model: A method for learning rumor propagation tree embedding based on hierarchical self-attention mechanism, which captures the propagation law by analyzing the cross-level interaction relationship between nodes.
[0043] UPFD Model: An end-to-end rumor detection method based on modeling user endogenous preferences and exogenous social backgrounds, which realizes detection by integrating individual behavior characteristics and social environment factors.
[0044] BiGCN Model: A detection method based on GCN, which enhances the modeling ability of information diffusion characteristics by learning the bidirectional information flow pattern of the rumor propagation tree.
[0045] GCAN Model: A detection method based on multi-source heterogeneous information fusion, which realizes detection through graph convolutional learning by integrating text semantics, user portraits, and network topology features.
[0046] DYNGCN Model: A detection method of dynamic graph convolutional network based on temporal snapshot partitioning, which conducts detection by learning the evolution of the propagation structure in stages.
[0047] DGNF Model: A detection method based on the GAT-Transformer hybrid architecture, which conducts detection by capturing long-range dependencies and local interactions in the propagation time series through the self-attention mechanism.
[0048] DECL Model: A rumor detection method based on contrastive learning of dynamic graphs with time snapshots, which captures the rumor evolution law by modeling the structural differences across time periods.
[0049] RDMSC Model: A joint detection method based on high-homogeneity social circle feature extraction and social interaction information fusion, which optimizes the discriminant performance by integrating multi-source heterogeneous features.
[0050] CoAHRD Model: A hybrid detection method based on content, context, user features, and co-attention mechanism, which realizes fine-grained classification through dynamic weight allocation.
[0051] These comparison models are all highly representative, covering different types of rumor detection methods, and can effectively compare with the method of the present invention in terms of performance, verifying the advantages and applicability in the rumor detection task from multiple dimensions.
[0052] 5. Experimental Results and Analysis The present invention will evaluate the performance of the proposed rumor detection framework from five aspects, including rumor comparison experiments, early detection capabilities, time complexity analysis, ablation experiments, and hyperparameter sensitivity analysis. First, by comparing with existing rumor detection models, the advantages of the method of the present invention in terms of overall detection performance, early rumor recognition ability, and computational efficiency are comprehensively analyzed. Secondly, ablation experiments are used to verify the roles of each module in the method of the present invention, and analyze their contributions to the final detection effect to prove their necessity. Finally, hyperparameter sensitivity analysis explores the impact of different hyperparameter configurations on the detection accuracy, and examines the adaptability of the method of the present invention under changes in key parameters.
[0053] (1) Results and analysis of comparison experiments Tables 2, 3, and 4 respectively show the experimental results of the method of the present invention on the Twitter15 dataset, Twitter16 dataset, and Pheme dataset. From the comparative analysis, it can be seen that the detection effects of the method of the present invention on these three datasets all exceed the existing methods, and the accuracy rates are increased by 1.4%, 1%, and 1.5% respectively. This shows that in the rumor detection task, fusing user structure information and joint modeling strategies can effectively improve performance. In addition, the experimental results further verify the adaptability and robustness of the method framework of the present invention in different task environments. The specific analysis is as follows: Traditional machine learning methods perform poorly. The experimental results show that the DTC model, RFC model, and SVM-TS model have poor detection performance, indicating that traditional methods based on manual features have limitations in identifying rumors. Traditional methods rely on artificial features, are easily affected by subjective biases, and are difficult to capture complex features in rumor propagation, resulting in low accuracy.
[0054] Deep learning methods are superior to traditional methods. The deep learning method (RvNN) has significantly better results in the detection task than traditional methods. This is mainly attributed to the ability to autonomously learn implicit patterns and complex features in the data, while traditional detection methods based on manual features often have difficulty comprehensively capturing this information, thus limiting the improvement of detection capabilities.
[0055] Advantages of user-text joint detection. Models such as the RDSML model and CoAHRD model that combine user behavior and text information have better detection performance than models that only rely on text information, indicating that integrating user social information and text content can more comprehensively capture rumor propagation characteristics. However, these methods do not fully consider the dynamic volatility in the rumor propagation process, especially the information-intensive characteristics in the outbreak interval and the impact of noise in the propagation process. In contrast, the method of the present invention can more effectively utilize these key information by focusing on mining the intensive features in the outbreak interval, thereby improving the accuracy of rumor detection.
[0056] Advantages of the method framework of the present invention. Whether in the four-classification task of the Twitter15 dataset and the Twitter16 dataset or the two-classification task of Pheme, the method of the present invention is significantly better than all comparison models. It takes into account the dynamic volatility in the rumor propagation process, mines dense features using the outbreak interval, and combines the joint features of users and comments, enabling it to capture various deceptive signals more comprehensively and accurately. Compared with traditional methods, the method of the present invention can better handle the complex temporal features and user interaction features in rumor propagation, thus improving the overall detection performance.
[0057] There are significant differences in the detection performance of the Twitter15 dataset, the Twitter16 dataset, and the Pheme dataset. Judging from the experimental results, the detection performance of the Pheme dataset is significantly lower than that of the Twitter15 dataset and the Twitter16 dataset, which can mainly be attributed to two key factors. First, the topic overlap is relatively high. The Pheme dataset only contains 5 breaking news events with highly overlapping content, causing the framework to be more inclined to identify topic categories rather than truthfulness. In contrast, the statements in the Twitter15 dataset and the Twitter16 dataset are more independent, with diverse information sources, enabling a more effective distinction between true and false information. Second, the available information is limited. The average number of posts for each statement in the Pheme dataset is only 26, far less than that of the Twitter15 dataset and the Twitter16 dataset, resulting in less available information during training, restricting the learning of the propagation pattern and context features by the method of the present invention and affecting the detection effect.
[0058] In summary, due to the problems of topic overlap and limited information volume in the Pheme dataset, the detection accuracy is relatively low, affecting the performance of the method of the present invention. While the Twitter15 dataset and the Twitter16 dataset have more independent topics and richer propagation information, enabling the method of the present invention to fully learn the propagation pattern of rumors and ultimately achieving better detection results. This indicates that in a diverse and information-rich data environment, the method of the present invention has stronger adaptability and detection capabilities.
[0059] Table 2 Rumor detection results of the Twitter15 dataset Table 3 Rumor detection results of the Twitter16 dataset Table 4 Pheme rumor detection results (2) Early rumor detection results and analysis Early detection of rumors is crucial for curbing the spread of rumors and mitigating their negative social impacts. Therefore, evaluating the performance of the method of the present invention at this stage has become a key indicator for measuring its effectiveness. To evaluate the early detection effect, the present invention compares different methods based on time, and the results are as Figure 3 and Figure 4 and Figure 5 shown. As can be seen from Figure 3 , as the time segmentation increases, the detection performance of each method gradually improves. It is worth noting that in the initial stage of rumor release, the accuracy of the method of the present invention is significantly better than that of other methods, and with the passage of time, its advantage becomes more prominent. This shows that the method of the present invention can effectively capture key features in the initial stage of dissemination and improve the detection accuracy. The main reason is that the method of the present invention makes full use of the dense features of users and texts for joint modeling, and can select representative feature intervals to avoid noise interference when more information is available, thereby improving the accuracy.
[0060] The method of the present invention is superior to other comparison methods in early detection, indicating that even in the case of less disseminated information, effective detection can still be relied on the background knowledge of source tweets. When the disseminated information is rich, the method of the present invention can combine background knowledge with disseminated information to further improve the detection accuracy. This fully proves the advantage of the method of the present invention in early rumor detection. In short, the method of the present invention has shown significant advantages in the early rumor detection task, can dynamically adjust strategies according to the dissemination environment, and significantly improve the detection accuracy.
[0061] (3) Results and analysis of time complexity To evaluate the computational time consumption of the method of the present invention, the present invention selects several relatively good comparison models (Bi-GCN model, DYNGCN model, GCAN model, DECL model, REMSC model, CoAHRD model) for comparison, and records the average inference time (unit: millisecond) for the model to predict the label of a single sample in each test set. The experimental results are shown in Table 5.
[0062] Table 5 Comparison of computational time results of the method of the present invention The experimental results show that while the method of the present invention improves the accuracy of rumor detection, although the inference time increases slightly, it still maintains reasonable computational efficiency. The main reason is that the method of the present invention enhances the ability to capture complex features at different propagation stages by jointly modeling user behavior and text information and combining a multi-level attention mechanism for feature fusion. However, this efficient information integration also brings an additional computational burden. To optimize the time complexity, the present invention uses dense features in the outbreak interval for detection during design instead of modeling the entire propagation process, thereby reducing the redundant computational cost. This optimization measure effectively reduces the computational complexity, making the inference time of the method of the present invention on large-scale datasets similar to that of other comparative models and maintaining strong application competitiveness.
[0063] In summary, although the inference time of the method of the present invention is slightly longer than that of other models on certain datasets, its significantly improved detection accuracy makes up for this shortcoming. After the optimized design, the method of the present invention is comparable to other methods in terms of computational efficiency and can provide efficient rumor detection services in practical applications, demonstrating its strong application potential.
[0064] (4)Results and analysis of ablation experiments To further analyze the effectiveness of each component in the method of the present invention, it was compared with its different variants. The specific ablation experiments include the following four parts: -U: Remove the user social network graph in method S1 and only use the post propagation network graph as the input, while keeping other parts unchanged. This experiment aims to verify the role of the user structure in the framework.
[0065] -A: Remove the global feature P part in method S3 and only retain the local features, aiming to verify the effectiveness of the global features in the framework.
[0066] -D: Remove the heterogeneous interaction part in method step S4, directly perform subsequent operations on the post propagation structure features and user structure features, and do not consider the fine-grained interaction features between the two, aiming to verify the contribution of fine-grained feature fusion.
[0067] -S: Remove the homogeneous interaction part in method step S5, simply splice the weighted heterogeneous interaction feature posts and the user structure ignoring the front-back association information and redundant noise in the temporal features, mainly to verify the contribution of selective fusion of temporal features to the overall framework.
[0068] The results of the ablation experiments are as Figure 6As shown, the performance changes of the method of the present invention on three datasets, Twitter15, Twitter16, and Pheme, are presented. The experiment aims to analyze the influence of each module on the detection effect to further verify the rationality of the design of the method of the present invention.
[0069] Regarding "-U", the experimental results show that after removing this module, the accuracy rates of the method of the present invention on the Twitter15, Twitter16, and Pheme datasets decrease by 5.6%, 5.4%, and 3.6% respectively. This indicates that the user structure is crucial for rumor detection. User behavior information provides an effective supplement in the absence of sufficient dissemination data, and the method of the present invention effectively improves the detection performance by jointly learning the dual-channel interaction features of posts and users. Therefore, introducing the user structure significantly enhances the robustness and accuracy of the method of the present invention.
[0070] Regarding "-A", the experimental results show that after removing this module, the accuracy rates of the method of the present invention decrease by 1.2%, 1.9%, and 1.5% respectively, verifying the necessity of global feature fusion. Global features provide richer context information, and working together with local features can improve the detection accuracy. Removing this module will result in the loss of some key context information, thus affecting the overall performance.
[0071] Regarding "-D", the experiment shows that after removing this module, the accuracy rates of the method of the present invention decrease by 2.2%, 2.6%, and 2.7% respectively. This result indicates that the heterogeneous interaction module plays a key role in improving the detection accuracy. By learning the fine-grained heterogeneous interaction between the post dissemination structure and the user structure features, it can effectively capture the complex relationship between the two, thereby enhancing the understanding and recognition ability of the rumor dissemination pattern. After lacking this module, only detecting by splicing information leads to a performance decline.
[0072] Regarding "-S", its removal causes the accuracy rates of the method of the present invention on the Twitter15, Twitter16, and Pheme datasets to decrease by 0.7%, 1.5%, and 0.6% respectively. The homogeneous interaction module can effectively reduce the interference of redundant and irrelevant information and optimize the temporal feature extraction by assigning weights through the self-attention mechanism. The experiment shows that the removal of this module weakens the model's selective learning ability for temporal information, thereby affecting the overall detection effect.
[0073] In summary, each module of the method of the present invention plays a key role in improving the overall performance. The user structure module provides additional behavioral information, the global feature fusion enhances the understanding of the context, the heterogeneous interaction module finely models the complex relationship between posts and users, and the homogeneous interaction module optimizes the extraction of temporal features and suppresses noise. The cooperation of each module enables the method of the present invention to perform excellently in the rumor detection task, with stronger generalization ability and stability.
[0074] (5)Experimental Parameter Results and Analysis To further optimize the performance of the method of the present invention, the present invention conducts hyperparameter analysis experiments, focusing on evaluating the influence of the following key parameters on the effect of the method of the present invention: Number of outbreak segment divisions (b): The present invention sets b to 1, 2, and 3 respectively, aiming to explore the influence of different numbers of outbreak segments on its detection effect. The division of outbreak segments is an identifier of key time periods in the rumor propagation process, and different division numbers may affect the ability of the method of the present invention to capture the propagation dynamics.
[0075] Proportion of deleted edges (q): This parameter controls the proportion of edges deleted in the graph, and is used to simulate noise and data incompleteness in information propagation. In the experiment, by adjusting the q value, the robustness under different noise conditions was tested.
[0076] Slope threshold (k): The slope threshold k is used to distinguish different outbreak interval ranges in the post propagation process. A reasonable slope threshold can more accurately identify the peak period of rumor propagation, and thus more precisely identify the key propagation time periods.
[0077] In this experiment, different settings were made for b, q, and k respectively, and evaluations were carried out on multiple datasets. According to Figure 7 、 Figure 8 、 Figure 9 The experimental results shown in, the present invention can draw the following conclusions: Influence of b: When b = 1, the detection accuracy of the method of the present invention is relatively low, probably because only relying on a single outbreak segment fails to fully capture the dynamic characteristics of rumor propagation. As b increases, when b = 2, the detection effect reaches the optimal, indicating that appropriate division helps to accurately depict the propagation dynamics, and two outbreak segments can cover most of the key information, thus improving the detection performance. However, when b = 3, although the detection accuracy is slightly improved, the increase is limited, and too many segments may introduce redundant information, increase the computational burden, and affect the detection efficiency. Considering both the detection accuracy and the computational cost, b = 2 is the optimal choice.
[0078] The influence of q: Appropriate edge deletion can remove noise and enhance robustness. Experiments show that the highest values are obtained when q is 0.2, 0.3, and 0.3 respectively, indicating that appropriate edge deletion helps to extract key information. However, an excessive edge deletion ratio will lead to information loss and reduce the detection accuracy. Therefore, a reasonable selection of q is the key to improving performance.
[0079] The influence of k: k affects the division of the outbreak interval and the detection performance. The experimental results show that the propagation peaks can be optimally identified when k is 0.4, 0.4, and 0.5. However, when the value of k is too high, the division criterion is more stringent, and only the steep propagation stage is retained. Although this helps to filter out weak propagation signals, it may lead to the loss of some key propagation patterns, thus affecting the detection accuracy. Therefore, k = 0.4 and k = 0.5 achieve a better balance between the ability to capture propagation patterns and information integrity, providing the best performance.
[0080] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A rumor detection method based on the fusion learning of dual-domain perception structure features, characterized in that, Including the following steps: S1. Obtain a social media message event dataset, construct a post propagation network graph and a user social network graph based on the dataset, and extract post propagation features and user social features respectively through a graph attention network encoder; S2. Use the propagation density peak and slope division method to identify the key time window of rumor diffusion, and extract relevant sub-structure features from the post propagation network graph and the user social network graph; S3. Process the post propagation sub-features and user social sub-features through a mutual attention mechanism to obtain post propagation fusion features and user social fusion features; S4. Construct a projection matrix through the post propagation fusion features and user social fusion features, and then use the projection matrix to perform weighted summation on the post propagation fusion features and user social fusion features to obtain post propagation interaction features and user social interaction features, and further generate the first weighted post propagation features and the first weighted user social features; S5. Perform homogeneous interaction information modeling on the first weighted post propagation features and the first weighted user social features to obtain the final homogeneous interaction comprehensive representation; S6. After splicing the final homogeneous interaction comprehensive representation with the original tweet features, input them into the rumor detection module for classification to obtain the rumor detection result.
2. The rumor detection method based on dual-domain perception structure feature fusion learning according to claim 1, wherein Step S1 specifically includes: The social media message event dataset , where represents the th event and is the total number of events; For each event Construct a post propagation network graph , where the root node of the propagation graph is the event declaration ; Let \(\mathcal{V}\) be the set of nodes in the post propagation network graph, and a node represents a tweet; Let \(\mathcal{E}\) be the set of edges in the post propagation network graph, representing the interaction between tweets; the adjacency matrix of the post propagation network graph is , when \(A_{ij} = 1\), it means that the \(i\)-th tweet has an interaction relationship with the \(j\)-th tweet; when \(A_{ij} = 0\), it means that the \(i\)-th tweet has no interaction relationship with the \(j\)-th tweet; For each event Construct a user social network graph , Let \(V\) be the set of nodes of the user social network graph, where a node represents a user; Let \(E\) be the set of edges of the user social network graph, representing the social relationships between users; the adjacency matrix of the user social network graph is , when \(a_{ij} = 1\), it means that the \(i\)-th user has a follow relationship with the \(j\)-th user; when \(a_{ij} = 0\), it means that the \(i\)-th user has no follow relationship with the \(j\)-th user; The post propagation network diagram and the user social network diagram are input into the Graph Attention Network (GAT) encoder for feature extraction to obtain post propagation features and user social features ; The GAT encoder adopts a two-layer GAT encoder. The first-layer GAT encoder preliminarily updates the node features, and the second-layer GAT encoder processes the output features of the first-layer GAT encoder.
3. The rumor detection method based on dual-domain perception structure feature fusion learning according to claim 2, wherein Step S2 specifically includes: For each event Calculate the propagation density , which is expressed by the following formula: , Among them, represents an indicator function, which takes the value of 1 when the time information belongs to the time , and 0 otherwise; the time is divided by hour; represents that the node belongs to the node set ; According to the event Based on the propagation time and propagation density, a propagation density map is obtained, and the local maximum value is detected in a sliding window manner to identify the outbreak nodes. The formula is as follows: , Among them, represents the time point of the local maximum, represents the peak value of the fluctuation during the propagation process, represents the window size, represents the operation of taking the maximum value; select the two peak values with the highest propagation density, denoted as the first peak time point , the second peak time point ; Calculate the slope within the neighborhood of each peak in the propagation density graph, and the formula is as follows: , Among them, represents the density value of the propagation density map at the time point , represents the window length; represents the slope, representing the change trend of the propagation density at this moment; From the local maximum point in the propagation density map corresponding to the peak value of the fluctuation during the propagation process Start, search for points with slopes of ± to both sides as the boundaries of the outbreak interval. Starting from the first peak time point scan left and right respectively to find the first time point that satisfies ; Determine the boundaries of the propagation outbreak interval: start time: ; end time: ; Finally, obtain the propagation outbreak interval corresponding to the first peak time point : ; Similarly, obtain the propagation outbreak interval corresponding to the second peak time point ; According to the obtained propagation outbreak interval, nodes and edges that meet the outbreak interval are screened out from the post propagation network diagram to obtain a post propagation subgraph , and the post propagation subgraph satisfies the following formula: , , Among them, represents the node set of the post propagation subgraph, represents the edge set of the post propagation subgraph; Through the above process, the first peak time point is extracted from the post propagation network graph of the post propagation sub-graph and the second peak time point of the post propagation sub-graph ; The post propagation sub-graph and are respectively subjected to feature extraction through the graph attention network GAT encoder to obtain the corresponding first post propagation sub-feature and the second post propagation sub-feature ; According to the association between the post propagation network graph and the user social network graph, the user social sub-graph corresponding to each post propagation sub-graph is extracted from the user social network graph, and the first user social sub-feature and the second user social sub-feature .
4. The rumor detection method based on dual-domain perception structure feature fusion learning according to claim 3, characterized in that Step S3 specifically includes: Propagate the first post feature , the second post propagation feature , the first user social feature and the second user social feature Perform feature mapping through a linear transformation function to obtain the first post propagation mapping feature , the second post propagation mapping feature , the first user social mapping feature and the second user social mapping feature ; The feature representation is fused using the multi-head attention mechanism. In the multi-head attention mechanism, each attention head learns a different attention distribution through independent parameters; for each attention head, the first post-propagation mapping sub-feature is calculated and the second post-propagation mapping sub-feature to obtain the attention weight matrix between them, which is expressed by the following formula: , Among them, represents the parameter of the th attention head, and Softmax represents the normalization function, represents the scaling factor; similarly, the first user social mapping sub-feature and the second user social mapping sub-feature are used to calculate the attention weight matrix ; Perform weighted summation on the first post propagation mapping sub-feature and the second post propagation mapping sub-feature to obtain the post propagation fusion sub-feature. The formula is as follows: , Among them, represents the post propagation fusion sub-feature output by the th attention head; similarly, the user social fusion sub-feature is obtained; in the multi-head attention mechanism, multiple attention heads calculate different attention distributions in parallel, and the outputs of multiple attention heads are concatenated together and mapped to the final feature space through a linear transformation. The formula is as follows , , Among them, represents the post propagation splicing sub-feature, represents the splicing operation, represents the first linear transformation matrix, represents the user social splicing sub-feature, represents the second linear transformation matrix; Post propagation splicing sub-feature and post propagation feature are fused to obtain the post propagation fusion feature, which is expressed by the following formula: , Among them, represents the post dissemination fusion feature; similarly, the user social fusion feature is obtained .
5. The rumor detection method based on dual-domain perception structure feature fusion learning according to claim 4, wherein Step S4 specifically includes: Propagate the fusion features through posts and the user social fusion features Construct a projection matrix , and then use the projection matrix to perform weighted summation on the post-propagated fusion features and the user social fusion features to obtain the post-propagation interaction features and the user social interaction features The formula is as follows: , , , Among them, represents the hyperbolic tangent activation function, represents the parameter of the projection matrix, represents the linear transformation matrix of the user's social features, represents the linear transformation matrix of the post propagation features; Respectively perform attention weight assignment on the post propagation interaction features and user social interaction features Calculate the importance of each feature dimension through the Softmax function, and perform weighted summation to obtain the first weighted post propagation feature and the first weighted user social feature : , , Among them, represents the Softmax function.
6. The rumor detection method based on dual-domain perception structure feature fusion learning according to claim 5, wherein, Step S5 specifically includes: Perform weight reallocation on the first weighted post propagation feature and the first weighted user social feature to learn the correlation and dynamic importance of internal features, and obtain the second weighted post propagation feature and the second weighted user social feature , which is expressed by the following formula: , , Among them, represents the self-attention mechanism; Use a gated recurrent unit (GRU) to perform temporal modeling on the features and to obtain the final hidden state as , and the formula is expressed as follows: , Among them, represents the operation of the gated recurrent unit; the first weighted post propagation feature , the first weighted user social feature and the final hidden state are concatenated to obtain the final homogeneous interaction comprehensive representation , which is expressed by the formula: .
7. The rumor detection method based on dual-domain perception structure feature fusion learning according to claim 6, characterized in that Step S6 specifically includes: Represent the final homogeneous interaction comprehensively with the event statement of the original tweet features to splice and obtain the final comprehensive features , which is expressed by the formula as follows: , Calculate the predicted probability vector through a fully connected layer and a Softmax function, and the formula is as follows: , Among them, represents a fully connected layer operation.
8. The rumor detection method based on dual-domain perception structure feature fusion learning according to claim 7, wherein, During the training process, optimize by minimizing the cross-entropy loss between the predicted probability and the true label distribution: , Among them, represents the cross-entropy loss, is the number of samples in the dataset, represents the number of classes, that is, the total number of classes in the classification task, indicates that this is the regularization term, representing the model parameters the sum of squares, is the regularization factor, represents the event the true label distribution.
Citation Information
Patent Citations
Social network rumor detection method based on emotion perception and graph convolutional network
CN116431760A
Collaborative attention network multi-mode rumor detection method fusing image features
CN118211122A
Rumor detection method based on dynamic heterogeneous graph attention network
CN118245861A
Social media rumor detection method based on graph convolutional network and social psychology
CN118585887A
Deep neural architectures for detecting false claims
US10803387B1
Cited By
Rumor detection method and system based on heterogeneous propagation structure and dynamic feature fusion
CN121456774A