Public opinion propagation prediction method based on dynamic time weighted Renyi entropy and graph neural network

By combining dynamic time-weighted Rényi entropy and graph neural network, a space-time fusion model is constructed, which solves the shortcomings of the public opinion propagation prediction model in the existing technology in terms of time dynamic characteristics and node influence assessment, and achieves more accurate public opinion propagation prediction and key node detection.

CN120493984APending Publication Date: 2025-08-15XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510582397.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing public opinion dissemination prediction model is difficult to fully consider the temporal dynamic characteristics and node influence assessment of network public opinion dissemination. Dynamic modeling is insufficient, and more attention is paid to local node characteristics, ignoring the quantitative role of dynamic entropy in propagation path prediction.

Method used

Using a method based on dynamic time-weighted Rényi entropy and graph neural network, a spatiotemporal fusion model is constructed, combined with node time-weighted Rényi entropy features and Node2Vec embedding features, space-time fusion model is carried out through GraphSAGE, and the complexity and propagation dynamics of network structure are captured, Rényi entropy indexes at two levels of local node entropy and global time step entropy are designed, a time-weighted mechanism is introduced, and the propagation uncertainty of network topology structure on different time nodes is quantified.

Benefits of technology

It improves the accuracy and interpretability of public opinion dissemination prediction, can more comprehensively portray the communication laws of online public opinion, enhances the performance of key node detection and transmission path prediction, and is suitable for prediction of different time points and public opinion dissemination environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493984A_ABST
    Figure CN120493984A_ABST
Patent Text Reader

Abstract

The invention discloses a public opinion propagation prediction method based on a dynamic time weighted Renyi entropy and a graph neural network, and the method comprises the steps: constructing a public opinion propagation network based on the hot topic information of a network platform; constructing a public opinion propagation prediction model based on the dynamic time weighted Renyi entropy and the graph neural network; and inputting the public opinion propagation network into the public opinion propagation prediction model to complete public opinion propagation prediction of the hot topic information. According to the method, the machine learning method, the complex network and the graph entropy theory are combined, a new view angle is provided for public opinion propagation prediction, a new thought is provided for application of the time weighted entropy features in complex network analysis, and the method has important theoretical value and practical significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of public opinion propagation prediction, and in particular to a public opinion propagation prediction method based on dynamic time-weighted Renyi entropy and graph neural network. Background Art

[0002] The scale-free nature of social networks and the cascading effects of information dissemination make them a core vehicle for the evolution of public opinion in modern society. As network size grows exponentially, the complexity of information dissemination pathways increases significantly. Studies have shown that over 78% of information dissemination on social networks follows a power-law distribution. This unbalanced propagation pattern allows public opinion surrounding emergencies to cascade across the entire network within hours. In public health emergencies, disinformation can spread up to six times more efficiently than true information. Because disinformation is generated and disseminated so rapidly, it can influence the emotions and behaviors of a large number of users in a short period of time. This can not only cause social panic and mislead public decision-making, but can also damage the reputation of governments and businesses, leading to economic and social instability.

[0003] Unlike traditional media dissemination, the decentralized nature of social networks makes every user a potential source of information, creating a fertile ground for the spread of false information. The underlying mechanism of the viral spread of rumors involves not only the heterogeneity of user behavior but is also closely related to the dynamic entropy changes in network structure and the evolution of community structure. In the field of public opinion dissemination modeling, existing methods have evolved in three stages: early studies mostly used differential equation models based on propagation dynamics, but their simplified assumptions about network structure limited prediction accuracy; mid-term graph network methods improved modeling capabilities by introducing community structure and node centrality, but still did not solve the problem of dynamic feature fusion; recent machine learning technologies have achieved breakthroughs through graph embedding and time series modeling, but still face bottlenecks in dealing with the causal relationships and heterogeneity of propagation paths.

[0004] In recent years, the application of graph entropy has expanded beyond measuring network complexity to predict information propagation paths and identify key nodes. In particular, in scenarios such as public opinion dissemination and viral transmission, graph entropy can help understand and predict interactions between nodes and the information dissemination process. As a measure of network complexity, graph entropy can characterize both global and local characteristics of a network from an information-theoretic perspective. Dehmer and Emmert-Streib studied the application of graph entropy in complex networks and proposed that graph entropy can reflect network complexity by measuring characteristics such as the network's topology and node distribution. While static network analysis based on Shannon entropy can characterize structural stability, it struggles to describe dynamic evolution. Renyi entropy, a generalization of information entropy, can be adapted to different types of network structures by adjusting the parameter α to control entropy sensitivity. However, its application in public opinion dissemination prediction has not been thoroughly explored.

[0005] Furthermore, many studies have begun leveraging machine learning, particularly graph-based deep learning methods, to simulate and predict information dissemination within networks, addressing the challenges traditional public opinion dissemination models face in navigating the complexity and dynamics of networks. Graph neural networks (GNNs) have been a significant breakthrough in recent years in analyzing network structured data. GraphSAGE (Graph Sample and Aggregation) stands out for its excellent inductive learning properties. It learns node representations by aggregating features of neighboring nodes, enabling GNNs to process large-scale graph data.

[0006] Key challenges facing current research on public opinion propagation prediction include: First, existing models struggle to fully account for the temporal dynamics of online public opinion dissemination, resulting in insufficient dynamic modeling. Traditional public opinion dissemination models are often based on static network assumptions, making it difficult to capture the time-varying nature of node interactions. While graph embedding techniques optimize node representations through random walk strategies, they still struggle to model temporal decay and causal relationships. Second, node influence assessments rely too heavily on static metrics such as degree centrality and focus primarily on local node characteristics, neglecting the quantitative role of dynamic entropy in predicting propagation paths. Summary of the Invention

[0007] To address these challenges, this study proposes a model for predicting the spread of online public opinion based on dynamic time-weighted Rényi entropy (DTWRE) and deep learning. This model integrates node-wise time-weighted Rényi entropy features with Node2Vec embedding features, creating an innovative spatiotemporal fusion model based on graph entropy theory and deep learning. Its core advantage lies in its ability to more comprehensively characterize the spread patterns of online public opinion by combining topological structure and propagation complexity. This combination not only improves the accuracy of public opinion propagation predictions but also takes into account temporal dynamics, making the predictions more interpretable and practical. This graph neural network-based public opinion propagation prediction model effectively captures the underlying patterns of public opinion propagation by combining network structural information, node characteristics, and historical propagation data. First, a dynamic Rényi entropy metric system is constructed based on generalized graph entropy theory, introducing a temporal weighting factor to jointly model network structural complexity and propagation dynamics. This entropy feature can flexibly capture entropy variations across different network structures and time windows by adjusting the entropy order α. Secondly, the node embedding process integrates the respective advantages of Rényi entropy features and graph embedding methods. By combining network structural information, node characteristics, and historical propagation data, a propagation prediction model is constructed that can automatically mine key information hidden in network data and effectively capture the dynamic changes in complex network structures. Finally, an evaluation system is established, combining metrics such as node Rényi entropy, time-step global Rényi entropy, and timeliness evaluation to verify the model's performance advantages in key node detection and propagation path prediction.

[0008] To achieve the above objectives, the present invention provides a method for predicting the spread of public opinion based on dynamic time-weighted Renyi entropy and graph neural network, comprising the following steps:

[0009] Build a public opinion dissemination network based on hot topic information on the network platform;

[0010] A public opinion propagation prediction model based on dynamic time-weighted Rényi entropy and graph neural networks is constructed. This model introduces a time-weighted mechanism and designs two-level Rényi entropy indicators: local node entropy and global time-step entropy. Furthermore, the DTWRE features are fused with the high-dimensional node embeddings generated by Node2Vec, and a spatiotemporal fusion modeling framework is constructed using GraphSAGE.

[0011] The public opinion propagation network is input into the public opinion propagation prediction model to complete the public opinion propagation prediction of hot topic information.

[0012] Preferably, Chinese rumor data, including a Chinese rumor data set of forwarded and commented information, is obtained from a false information reporting platform of a social networking software and preprocessed; based on the preprocessed data, the public opinion propagation network is constructed; the preprocessing method includes: unifying time tags, converting to timestamp format, performing data enhancement processing, removing isolated nodes, and dividing time windows.

[0013] Preferably, the expression of the local node entropy includes:

[0014]

[0015] Among them: H a (v,t) represents the local node entropy; N(v,t) represents the neighbor set of node v at time step t; represents the normalized probability distribution of the information metric of neighbor node u; α is the order parameter of the Rényi entropy, which is used to adjust the sensitivity to different distribution sparsity.

[0016] Preferably, the expression of the global time step entropy includes:

[0017]

[0018] in, is time t k The entropy of the network snapshot at a given moment; V t Represents the set of nodes in a snapshot; represents the global time step entropy; ω(tt k ) is the time weight function, using exponential decay weight:

[0019]

[0020] Where λ>0 controls the contribution of different time steps to the final entropy value.

[0021] Preferably, the GraphSAGE workflow includes:

[0022] Neighbor sampling: randomly sampling a fixed number of neighbors of each node;

[0023] Feature aggregation: For each node u, GraphSAGE aggregates features from neighboring nodes and updates its own representation as follows:

[0024]

[0025] in, is the feature representation of node u at the kth layer; Aggregate{} is the aggregation function; W k is a trainable parameter; σ is a nonlinear activation function; N represents the set of neighboring points; Represents the feature representation of node v at the k-1 layer;

[0026] Final representation: After multiple layers of aggregation, the final representation of node v is:

[0027]

[0028] in, Represents the feature representation of node (v, d) at the nth layer.

[0029] Preferably, the public opinion propagation prediction model includes two parts: a feature input layer and a graph neural network layer;

[0030] The feature input layer is used to use the extracted node features as the input of the model;

[0031] The graph neural network layer is used to update the representation of the node by aggregating the features of neighboring nodes through GraphSAGE, and generate embedded representations of the node in different time windows.

[0032] Preferably, the public opinion propagation prediction model uses a binary cross entropy loss function to train the link probability of the node pairs output by the model:

[0033]

[0034] Among them, y uv Indicates whether the node pair (u,v) is connected, is the link probability predicted by the model; E represents the set of real edges in the network.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] This invention combines machine learning methods, complex networks and graph entropy theory, which not only provides a new perspective for public opinion propagation prediction, but also provides new ideas for the application of time-weighted entropy features in complex network analysis. It has important theoretical value and practical significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 A schematic diagram of time window division according to an embodiment of the present invention;

[0039] Figure 2 This is a schematic diagram of the spatiotemporal fusion modeling process according to an embodiment of the present invention;

[0040] Figure 3 Schematic diagram of the workflow of GraphSAGE according to an embodiment of the present invention.

[0041] Figure 4 Schematic diagram of the experimental process of an embodiment of the present invention;

[0042] Figure 5 Schematic diagram of experimental comparison results of an embodiment of the present invention;

[0043] Figure 6 Schematic diagram of the effect of the DTWRE order α on the model in an embodiment of the present invention;

[0044] Figure 7 Schematic diagram of the effect of the time weight parameter λ on the model in an embodiment of the present invention;

[0045] Figure 8 Schematic diagram of the effect of the ratio of positive and negative samples on the model in an embodiment of the present invention;

[0046] Figure 9 Schematic diagram of timeliness evaluation results in an embodiment of the present invention;

[0047] Figure 10 This is how the indicators of the real public opinion dataset change with the training epoch during the model operation in the embodiment of the present invention;

[0048] Figure 11 Schematic diagram of a social public opinion network in an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0050] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] Example 1

[0052] This embodiment provides a method for predicting the spread of public opinion based on dynamic time-weighted Renyi entropy and graph neural network, including the following steps:

[0053] S1. Build a public opinion dissemination network based on hot topic information on the network platform.

[0054] This example crawls Chinese rumor data from Sina Weibo's false information reporting platform, including a dataset of forwarded and commented Chinese rumor messages. Predicting this rumor data can reveal the propagation characteristics of false public opinion, thereby providing theoretical support for the governance of false information on social networks.

[0055] The following data preprocessing is required for the above original data set:

[0056] Unify time labels and convert them into timestamp format.

[0057] Considering the scale of real social network datasets, data enhancement is performed.

[0058] Removing isolated nodes: Isolated nodes in a social network do not participate in propagation, so all isolated nodes need to be removed to ensure the connectivity of the network graph. This is also to ensure the rationality of calculating the Rényi entropy.

[0059] Time window division: In order to effectively capture the dynamic changes of online public opinion, this embodiment divides the propagation data into multiple time windows according to the timestamp. Each time window corresponds to a time step, and the weight gradually changes over time. Different data sets are divided into different time steps based on experience. The division diagram is as follows Figure 1 shown.

[0060] The real social network dataset contains information such as the content of rumor microblogs, microblog publishers, user forwarding and comment records, and interaction time. This embodiment needs to filter out the interactive relationship between rumor publishers and users' comments and forwarding of the microblogs to construct a public opinion dissemination network.

[0061] S2. Construct a public opinion propagation prediction model based on dynamic time-weighted Rényi entropy and graph neural network.

[0062] The public opinion propagation prediction model of this embodiment introduces a time-weighted mechanism and designs a two-level Rényi entropy metric: local node entropy and global time-step entropy. This effectively quantifies the uncertainty and complexity of network topology propagation at different time nodes. Furthermore, by integrating DTWRE features with high-dimensional node embeddings generated by Node2Vec and constructing a spatiotemporal fusion modeling framework using GraphSAGE, this model achieves accurate predictions of link generation and key node identification in public opinion propagation.

[0063] The core idea of the network public opinion propagation prediction model based on DTWRE combined with GraphSAGE is to use time-weighted Rényi entropy to quantify the complexity changes of the network structure at different time nodes in the public opinion propagation network. At the same time, this entropy feature is combined with the node feature vector obtained by graph embedding technology, and then transmitted to the downstream graph neural network for link prediction tasks, ultimately realizing public opinion propagation prediction.

[0064] Among existing methods for predicting the spread of online public opinion, Node2Vec uses random walks to learn network topology information. This method can capture local and global node associations and the topological characteristics of propagation paths. The resulting low-dimensional vectors retain a richer representation of the network structure. However, they cannot directly depict the temporal evolution of public opinion propagation, lacking the ability to model temporal dynamics and unable to distinguish between recently formed and historical connections. Furthermore, the Node2Vec algorithm has another inherent drawback: its vector representation is data-driven, making it difficult to directly interpret its physical meaning, such as node importance and propagation capacity. Incorporating the Renyi entropy feature effectively overcomes this shortcoming.

[0065] This embodiment's innovative time-weighted Rényi entropy measures the uncertainty of node propagation, introduces a time dimension, increases the model's predictive depth, and captures the evolutionary trends of public opinion dissemination. Rényi entropy can measure the uncertainty of node propagation, and the designed time-weighted mechanism more accurately reflects the dynamic complexity of nodes in the public opinion dissemination process. It also models the importance of different time points during the dissemination process, enabling the model to distinguish key nodes such as early influencers and late disseminators, thus enhancing research on public opinion dissemination in the spatiotemporal dimensions.

[0066] However, a single entropy value is difficult to fully express the characteristics of the network structure, resulting in less information than the high-dimensional embedding of Node2Vec. Therefore, this embodiment combines the complementary characteristics of the two. After the combination, Node2Vec can learn the global network structure and capture the topological characteristics of the propagation path. The time-weighted Rényi entropy provides temporal dynamic information, enabling the model to perceive the changing laws of public opinion propagation over time. GraphSAGE can simultaneously learn spatial topological relationships and temporal evolution trends, which is more in line with the propagation model of real social networks. The entire spatiotemporal fusion modeling process is as follows: Figure 2 As shown in the figure. By leveraging the high-dimensional topological information provided by Node2vec and the temporal dynamic characteristics of Rényi entropy, and training with GraphSAGE, we fully integrate network structure and propagation characteristics, thereby improving the accuracy, generalization ability, and interpretability of public opinion propagation predictions, enhancing the stability of the model, and providing a more performant prediction model for social network public opinion analysis.

[0067] Thanks to the entropy order α of the Rényi entropy, the sensitivity to different network structures can be adjusted, and the entropy value can be calculated at the local node level and the global network level, making the model more suitable for predictions at different time points and in different public opinion dissemination environments (such as emergencies, slow-heating topics, etc.).

[0068] In this example, based on the definition of Renyi entropy, we designed and applied two hierarchical entropy metrics: local node entropy (LNE) and global time-step entropy (DTWRE). These metrics capture the information complexity at different scales during the spread of online public opinion, and serve as features for subsequent GraphSAGE model learning. DTWRE comprehensively considers the propagation characteristics of individual nodes and the overall network complexity, which incorporates the influence of propagation history on the current state. This provides an effective metric for modeling the complexity of public opinion dissemination.

[0069] Rényi entropy is a generalized entropy measure used to measure the uncertainty of network structure. The mathematical definition of LNE is as follows:

[0070]

[0071] Among them: H a (v,t) represents the local node entropy; N(v,t) represents the neighbor set of node v at time step t; represents the normalized probability distribution of the information metric (such as node degree or communication influence) of neighbor node u; α is the order parameter of the Rényi entropy, which is used to adjust the sensitivity to different distribution sparsity.

[0072] To improve the model's predictive performance, this embodiment introduces a time factor and uses a time-weighted function to perform a weighted summation of the entropy values of different time snapshots. This allows more recent information to have a greater impact on the entropy value, thereby closely simulating the mechanism of real public opinion dissemination. The calculation formula for the created DTWRE is as follows:

[0073]

[0074] in, is time t k The entropy of the network snapshot at a given moment; V t Represents the set of nodes in a snapshot; represents the global time step entropy; ω(tt k ) is the time weight function, using exponential decay weight:

[0075]

[0076] Where λ > 0 controls the contribution of different time steps to the final entropy value. Larger λ values cause the model to focus more on recently disseminated information, while smaller λ values allow the model to consider public opinion spread over a longer timeframe. LNE and DTWRE will subsequently integrate Node2Vec embedding features (as shown in Table 1) to provide feature input for downstream learning models.

[0077] Table 1

[0078]

[0079] In order to use graph structure information to predict the spread of public opinion, this embodiment uses GraphSAGE to perform public opinion prediction tasks. Among the current mainstream graph neural networks, graph convolutional networks (GCN) and graph sampling aggregation networks (GraphSAGE) are two very important models. Both GCN and GraphSAGE are committed to learning useful patterns and features from graph structure data. Based on a large amount of graph data for training, they learn the parameters of the model to adapt to different graph data and make accurate predictions. By processing and analyzing the information of nodes and edges, they can achieve tasks such as node classification and link prediction.

[0080] GCN is a method that combines topological structure and vertex attribute information to learn vertex embedding representations in a graph. However, GCN requires learning vertex embeddings within a fixed graph and cannot directly generalize to vertices that have not appeared during training. This is a form of transductive learning. This means that when new nodes appear in the graph, the model must be retrained, which is computationally expensive and unsuitable for dynamically changing or large-scale graph data.

[0081] GraphSAGE, introduced in this example, is an inductive learning framework that can efficiently generate unknown vertex embeddings using vertex attribute information. Its core idea is to generate the target vertex embedding vector by learning a function that aggregates neighboring vertices. The specific workflow is as follows: Figure 3 Its main features include:

[0082] 1. Neighbor Sampling: Since social networks are usually large-scale sparse graphs, GraphSAGE randomly samples a fixed number of neighbors for each node to reduce computational complexity. This is very important for large-scale graph data computation.

[0083] 2. Feature aggregation: For each node u, GraphSAGE aggregates features from neighboring nodes and updates its own representation as follows:

[0084]

[0085] in, is the feature representation of node u at the kth layer; Aggregate{} is the aggregation function, where mean aggregation is used; W k is a trainable parameter; σ is a nonlinear activation function, ReLU is used here; N represents the set of neighboring points; Represents the feature representation of node v at the k-1th layer.

[0086] 3. Final representation: After multiple layers of aggregation, the final representation of node v is:

[0087]

[0088] in, Represents the feature representation of node (v, d) at the nth layer.

[0089] In this embodiment, DTWRE is added as an additional feature to the input of GraphSAGE, so that the model can learn the temporal dynamic characteristics of the propagation process.

[0090] S3. Input the public opinion dissemination network into the public opinion dissemination prediction model to complete the public opinion dissemination prediction of hot topic information.

[0091] The constructed public opinion propagation network is connected to the constructed public opinion propagation prediction model to complete the prediction of public opinion propagation.

[0092] Example 2

[0093] This example is intended to demonstrate the experimental process and results of the model of the present invention on real data. The experimental platform is Pycharm, using development tools such as PyTorch and DGL (Deep Graph Library). DGL is a deep learning framework for graph neural networks (GNNs), which provides efficient, flexible, and easy-to-use tools for the research and development of graph neural networks. In order to fully verify the performance of the model of the present invention, this example will compare it with other existing link prediction methods to demonstrate its excellent performance on different evaluation indicators.

[0094] To evaluate the model proposed in this invention, this example used the following datasets in the verification and empirical phases:

[0095] CollegeMSG Dataset: This classic dataset from SNAP consists of private messages sent on an online social network. As a temporal network, this dataset is common in network research. This study used it during model validation to facilitate performance comparisons with other research methods.

[0096] Real social network dataset: as step S1 in the first embodiment.

[0097] After the data collection is completed, it is pre-processed, and the steps are similar to S1 in the first embodiment.

[0098] The goal of the public opinion prediction task in this example is to predict possible future links based on a known network structure. The ratio of positive and negative samples not only affects the training and learning of the model, but also its generalization ability. The positive and negative samples constructed are as follows:

[0099] Positive samples: connected node pairs (u,v) from the real network.

[0100] Negative samples: Using negative sampling strategy, randomly sample unconnected node pairs (u, v), that is Where E represents the set of real edges in the network.

[0101] However, real networks are often sparse, meaning that most possible edges do not exist. This raises the issue of negative sample selection. Improper negative sampling can lead to data imbalance and reduce the model's discriminative ability. To avoid this problem, this implementation oversamples negative samples.

[0102] In this example, to fully capture the dynamic information and structural characteristics of online public opinion propagation, we constructed multidimensional features as input to the GraphSAGE model. The main features used in this example can be divided into two categories: information entropy-based features and node embedding features.

[0103] (1) Entropy characteristics

[0104] As previously explained, this approach encompasses both local and global features. LNE quantifies the propagation potential and local structural complexity of each node at a specific time step. For any node v in the network, its LNE at time step t serves as its own characteristic. DTWRE measures the structural complexity of the entire network at a specific time step. Combined, these two metrics comprehensively measure the complexity of public opinion social networks at different stages of propagation.

[0105] (2) Embedded features

[0106] To further capture the semantic and structural information between nodes, this embodiment uses the Node2Vec algorithm to generate node embeddings. The specific steps are as follows:

[0107] 1. Graph conversion: Convert the graph built by DGL to NetworkX format and convert it into an undirected graph to meet the requirements of Node2Vec.

[0108] 2. Embedding calculation: Use the Node2Vec model to perform random walk sampling on the network and generate continuous vector representations of nodes through the Skip-Gram model. The embedding dimension is set to 64.

[0109] 3. Standardization: To eliminate scale differences, the generated node embeddings are normalized using MinMaxScaler and converted to Tensor format to ensure consistency with other features.

[0110] Node embedding features can complement the deficiencies of information entropy features in capturing local structures and semantic relationships, thereby enabling the model to obtain richer node representations in link prediction tasks.

[0111] (3) Feature Fusion

[0112] In the final feature construction stage, this embodiment concatenates LNE, DTWRE, and Node2Vec embeddings to form a comprehensive feature vector.

[0113] This feature fusion approach combines information at multiple scales. LNE captures the uncertainty of a single node's local structure, DTWRE reflects network complexity across the entire time step, and node embedding provides high-dimensional semantic information between nodes. Furthermore, through a time-weighted strategy, the model can flexibly reflect the impact of historical information on the current state, helping to capture the temporal nature of public opinion dissemination. The combined features provide a more comprehensive input for the GraphSAGE model, enabling more accurate discrimination between positive and negative examples in link prediction tasks, improving prediction accuracy.

[0114] The specific process of model construction and training is as follows:

[0115] 1. Model Architecture: The model in this embodiment consists of two main parts:

[0116] (1) Feature input layer: This layer uses the extracted node features (entropy features and embedding features) as the input of the model.

[0117] (2) Graph neural network layer: It updates the representation of nodes by aggregating the features of neighboring nodes through GraphSAGE and generates embedded representations of nodes in different time windows.

[0118] 2. Predictor and loss function: To train the model for effective link prediction, the predictor uses an MLP multi-layer perceptron, which consists of multiple fully connected layers and can learn the nonlinear relationship between node features. The predictor ultimately generates a prediction score that represents the likelihood of a link. The loss function uses a binary cross-entropy loss function, which is used to train the link probability of the node pair output by the model. Specifically, given two node pairs (u, v), the model needs to predict whether the two nodes will form a connection within a certain time window in the future. The specific calculation formula is:

[0119]

[0120] Among them, y uv Indicates whether the node pair (u,v) is connected, is the link probability predicted by the model.

[0121] 3. Training Process: During training, we used the Adam optimizer, which offers advantages such as adaptive learning rates and fast convergence. We also used cross-validation to adjust hyperparameters. The model was trained over 100 batches to minimize the loss function and achieve optimal training results.

[0122] To facilitate comparison with other research methods and comprehensively evaluate the performance of the model proposed in this invention, this example uses the following most common evaluation indicators:

[0123] AUC (Area Under the ROC Curve): The AUC value measures the model's ability to distinguish between positive and negative samples. A larger value indicates a stronger predictive ability of the model.

[0124] Precision: Precision indicates how many of the pairs of nodes predicted as positive are actually positive. The higher the precision, the better the prediction accuracy of the model.

[0125] Recall: Recall indicates how many node pairs are actually positive that the model successfully predicts as positive. A higher recall indicates a better coverage of positive node pairs by the model.

[0126] F1-Score: The F1 score is the harmonic mean of precision and recall, which comprehensively considers the predictive ability of the model.

[0127] Temporal Performance Evaluation: Since public opinion propagation has obvious temporal evolution characteristics, we evaluate the performance of the model at different time steps, focusing on the changes in the model's accuracy over different time periods.

[0128] In summary, the key steps of the entire experiment include data processing, feature engineering, model training and evaluation, such as Figure 4 shown.

[0129] During the model training and evaluation process, we conducted a lot of experimental settings and hyperparameter adjustments. During model training, some parameters in the model need to be continuously tuned during the experiment to obtain the optimal value and improve model performance:

[0130] 1. Time window size: This affects the model's ability to learn temporal information. It is necessary to set different window sizes based on experience for different data sets to better capture the dynamics of public opinion dissemination.

[0131] 2. α value (entropy order of LNE): This determines the flexibility of entropy calculation. We tune α using a grid search method to select the optimal value.

[0132] 3. λ value (time weight parameter): In the experiment, the λ value is tuned based on different time steps and historical experience of public opinion dissemination.

[0133] 4. Negative Sampling Ratio: During training, the model needs to sample a certain proportion of negative samples. We set different negative sampling ratios to observe their impact on model performance.

[0134] To verify the effectiveness of the proposed method, this example designed multiple baseline models in the experiment and compared the results with the following classic prediction methods:

[0135] ① Traditional node degree-based method;

[0136] ② Traditional method based on node PageRank value;

[0137] ③Node2Vec method based on node embedding;

[0138] ④Based on the traditional static Rényi entropy method;

[0139] These methods can fully verify the innovativeness of the method of the present invention, and reasonably compare the contribution of different features to the model performance.

[0140] Example 3

[0141] This example demonstrates the experimental results of a dataset using the proposed online public opinion prediction model based on DTWRE and deep learning. Through comparative experiments, performance evaluation, timeliness analysis, and real-world public opinion data expansion, we validated the effectiveness of the proposed method in online public opinion prediction tasks and analyzed the impact of various factors on model performance. The experimental results demonstrate that the graph neural network model combined with DTWRE achieves higher accuracy and greater timeliness than traditional methods.

[0142] The experiment compared the four baseline methods mentioned above. The evaluation index comparison results are shown in Table 2:

[0143] Table 2

[0144]

[0145] According to Table 2 and Figure 5 Overall, our innovative method, "DTWRE," achieved the highest AUC (0.9742), demonstrating that the dynamic time weighting mechanism better captures the underlying dynamic propagation characteristics of the network when distinguishing between positive and negative examples, thereby improving the model's discriminative ability. The DTWRE method also achieved the highest Precision, 0.9259, indicating that the vast majority of predicted positive examples are real links. This result demonstrates the effectiveness of node embeddings in capturing node semantic and structural information, and demonstrates that relying solely on embeddings may not fully reflect the dynamic nature of public opinion dissemination. In terms of recall, both DTWRE (0.9144) and Rényi entropy (0.8970) outperformed methods based on node degree (0.8618) and PageRank (0.8613), demonstrating that the inclusion of entropy features more comprehensively captures real links, thereby reducing the rate of missed detections. The innovative method achieved the highest F1-Score (0.9201), indicating a good balance between precision and recall. Compared to Node2Vec's F1-Score (0.9062), the temporal weighting mechanism demonstrates its superiority in overall performance. Regarding accuracy, DTWRE achieved an accuracy of 0.9207, still the highest among all methods, demonstrating its stability in overall prediction tasks.

[0146] In summary, since information at different time steps has different effects on the network structure during the process of public opinion propagation, the Rényi entropy can provide nodes with richer propagation information, thereby helping the model better capture the potential characteristics of the network structure. In addition, the time-weighted entropy feature can adapt to the time-varying public opinion propagation characteristics in the network, making the model perform better when dealing with dynamic propagation processes. The DTWRE method proposed in this paper effectively integrates the information within each time step by introducing exponential decay weights, and ultimately surpasses all baseline methods in all indicators, proving the effectiveness and advancement of this method in the task of predicting the spread of online public opinion.

[0147] Example 4

[0148] To better illustrate the impact of key parameters in the method of the present invention on link prediction performance, this example conducts a large number of experiments on the DTWRE order α, the time weight factor λ, and the ratio of positive and negative samples. The impact of changes in the values of each parameter on various model indicators is demonstrated through visual charts.

[0149] (1) DTWRE order α

[0150] By changing the DTWRE order α, we explore its impact on the model performance. The results show that Figure 6 As shown in the figure, when α is 0.2, the AUC is 0.950, indicating that the entropy calculation has a low sensitivity and may not fully capture the structural diversity between node neighbors. As α increases to 0.6, the AUC increases significantly to 0.966, indicating that the entropy calculation can more effectively reflect the distribution uncertainty between node neighbors, thereby improving the model's ability to distinguish between positive and negative samples. When α continues to increase to 1, 1.5, 2, and 5, the AUC decreases to 0.959, 0.955, 0.952, and 0.944, respectively. This indicates that an excessively high α will overemphasize the influence of larger values in the probability distribution and ignore the information of the lower probability part, resulting in a decrease in overall discrimination ability. Precision reaches its highest value (0.922) at α = 0.6 and then decreases slightly as α increases. The trends of Recall and F1-Score are similar to Precision, and are optimal when α = 0.6. Neither too low nor too high α can achieve the optimal balance. This change indicates that when α is set to 0.6, entropy can better capture the diversity of node propagation potential, making the model more accurate in distinguishing positive and negative samples. However, setting it too low or too high can lead to information loss or the introduction of noise, thus affecting overall performance. The trend of Accuracy is similar to that of AUC, rising from 0.888 (α = 0.2) to 0.916 (α = 0.6) and then gradually decreasing to 0.904 (α = 5), further proving that α = 0.6 is the optimal value.

[0151] Therefore, the modified data shows that all indicators reach their optimal state when α = 0.6, which is consistent with our expectation of the DTWRE feature's strength in capturing the dynamics of public opinion dissemination. This data change highlights the unique role of entropy in revealing node influence and propagation uncertainty. Compared to traditional metrics that only consider node degree, it can more comprehensively describe the complexity of public opinion diffusion.

[0152] (2) Influence of the time weight parameter λ

[0153] The parameter λ in the time weighting factor is an important innovation in the DTWRE of this invention. It is necessary to explore its influence on the model. The results are as follows: Figure 7 shown.

[0154] It can be seen that when the value of λ is low (such as 0.1), the overall performance of the model is poor due to insufficient weight of historical information; as λ increases from 0.1 to a medium value (0.4 to 1.2), the performance indicators of the model are significantly improved, indicating that moderately increasing the weight of recent information helps to capture the characteristics of dynamic public opinion dissemination; when λ reaches 1.2, the model achieves excellent levels in AUC, Recall, F1-Score and Accuracy, and has the best overall performance; when λ is further increased to 2, although Precision and Accuracy increase slightly, the recall rate decreases slightly, and the overall indicators do not change much, indicating that a large λ may lead to excessive neglect of historical information.

[0155] (3) Impact of the ratio of positive and negative samples

[0156] In the present invention, the ratio of positive and negative samples needs to be carefully considered, because when the ratio of positive and negative samples is seriously unbalanced, the model may tend to predict the sample type with a larger number, and a suitable ratio of positive and negative samples helps the model better learn the features and patterns in the data, thereby improving the performance and prediction accuracy of the model. Figure 8 The effect of different positive and negative sample ratios on performance indicators is shown.

[0157] As can be seen, the ratio of positive and negative samples significantly affects model performance. When the ratio is set to 2 (too many negative samples), while precision is high, recall drops significantly, resulting in low overall F1-Score and Accuracy. When the ratio is set to 0.5 (fewer negative samples), recall is high but precision is insufficient. The optimal positive-to-negative sample ratio is around 0.75, at which the model achieves the best performance in terms of AUC, F1-Score, and Accuracy, indicating that this ratio achieves the best balance between positive and negative samples.

[0158] (4) Timeliness evaluation

[0159] To verify the present invention's ability to capture the temporal dynamics of online public opinion dissemination, this example designed experiments with different time steps. Specifically, while maintaining the same parameters such as λ, α, and the ratio of positive and negative samples, the impact of different time steps on model performance was compared. The experimental results in Table 3 show performance indicators at three different time steps: 1.75 days, half a week, and a week.

[0160] Table 3

[0161]

[0162] from Figure 9 As can be seen, using a longer time window length, when appropriate for the total time span of the dataset, offers significant advantages. When a time window length of 604,800 seconds (7 days) is used, all metrics significantly outperform shorter time window lengths (302,400 seconds and 151,200 seconds). This demonstrates that a 7-day time window length fully captures the dynamics of public opinion dissemination and provides sufficient information for predicting link generation. Shorter time windows, on the other hand, may struggle to accurately predict due to insufficient data, leading to lower overall performance. It's worth noting that the 7-day time window length was designed based on empirical experience and the total time span of the dataset, so only three suitable time window lengths are explored here. Therefore, from a timeliness perspective, using a longer time window length is more effective in ensuring the model's robust capture of network dissemination dynamics.

[0163] Example 5

[0164] In the above examples, we can see the excellent performance of the model on the dataset. In order to verify the generalization ability of the model, this example further analyzes the spread of public opinion on a real Weibo dataset. This dataset covers interactive information such as forwarding and commenting of rumors among Weibo users. In order to explore the propagation mechanism of false public opinion, relying on the advantages of the innovative DTWRE of this study, the results of the rumor data analysis are as follows: Figure 10 shown.

[0165] Figure 10 This chart shows how the metrics of a real public opinion dataset change over the training epochs. Overall, the curve shows a significant initial rapid rise, then gradually converges and stabilizes around the 50th epoch, reaching a relatively high level. This demonstrates excellent overall performance, demonstrating the model's robust ability to comprehensively discern positive and negative links (i.e., public opinion transmission relationships) in real public opinion datasets.

[0166] Figure 11The entire social public opinion network exhibits a "center-periphery" structure. Red nodes distributed in the central area have higher entropy values and are considered "key nodes." They form a dense core of communication at the center of the network, rapidly disseminating information to ordinary nodes on the periphery during the public opinion propagation process, and therefore require special attention in public opinion governance. The importance of these nodes can be verified by reverse engineering the nodes in the original dataset and analyzing their propagation paths. Red lines represent links predicted by the model as "positive," indicating an edge with a higher likelihood of actual connection or information flow during the public opinion propagation process. Red lines can be seen radiating from the center to the periphery, indicating that the model confirms the information propagation paths from key nodes to ordinary nodes.

[0167] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A method for predicting public opinion propagation based on dynamic time-weighted Renyi entropy and graph neural network, characterized by the following steps: include: Build a public opinion dissemination network based on hot topic information on the network platform; A public opinion propagation prediction model based on dynamic time-weighted Rényi entropy and graph neural networks is constructed. This model introduces a time-weighted mechanism and designs two-level Rényi entropy indicators: local node entropy and global time-step entropy. Furthermore, the DTWRE features are fused with the high-dimensional node embeddings generated by Node2Vec, and a spatiotemporal fusion modeling framework is constructed using GraphSAGE. The public opinion propagation network is input into the public opinion propagation prediction model to complete the public opinion propagation prediction of hot topic information.

2. The method for public opinion propagation prediction based on dynamic time-weighted Renyi entropy and graph neural network according to claim 1 is characterized in that: Obtain Chinese rumor data from the false information reporting platform of social networking software, including a Chinese rumor dataset of forwarded and commented information, and perform preprocessing; Based on the pre-processed data, construct the public opinion dissemination network; The preprocessing method includes: unifying time tags, converting to a timestamp format, performing data enhancement processing, removing isolated nodes, and time window division.

3. The method for public opinion propagation prediction based on dynamic time-weighted Renyi entropy and graph neural network according to claim 1 is characterized in that: The expression of the local node entropy includes: Among them: H α (v, t) represents the local node entropy; N(v, t) represents the neighbor set of node v at time step t; represents the normalized probability distribution of the information metric of neighbor node u; α is the order parameter of the Rényi entropy, which is used to adjust the sensitivity to different distribution sparsity.

4. The method for public opinion propagation prediction based on dynamic time-weighted Renyi entropy and graph neural network according to claim 3 is characterized in that: The expression of the global time step entropy includes: in, is time t k The entropy of the network snapshot at a given moment; V t Represents the set of nodes in a snapshot; represents the global time step entropy; ω(tt k ) is the time weight function, using exponential decay weight: Where λ>0 controls the contribution of different time steps to the final entropy value.

5. The method for public opinion propagation prediction based on dynamic time-weighted Renyi entropy and graph neural network according to claim 1 is characterized in that: The GraphSAGE workflow includes: Neighbor sampling: randomly sampling a fixed number of neighbors of each node; Feature aggregation: For each node u, GraphSAGE aggregates features from neighboring nodes and updates its own representation as follows: in, is the feature representation of node u at the kth layer; Aggregate{} is the aggregation function; W k is a trainable parameter; σ is a nonlinear activation function; N represents the set of neighboring points; Represents the feature representation of node v at the k-1 layer; Final representation: After multiple layers of aggregation, the final representation of node v is: in, Represents the feature representation of node (v, d) at the nth layer.

6. The method for public opinion propagation prediction based on dynamic time-weighted Renyi entropy and graph neural network according to claim 1 is characterized in that: The public opinion propagation prediction model consists of two parts: a feature input layer and a graph neural network layer; The feature input layer is used to use the extracted node features as the input of the model; The graph neural network layer is used to update the representation of the node by aggregating the features of neighboring nodes through GraphSAGE, and generate embedded representations of the node in different time windows.

7. The method for public opinion propagation prediction based on dynamic time-weighted Renyi entropy and graph neural network according to claim 6 is characterized in that: The public opinion propagation prediction model uses a binary cross entropy loss function to train the link probability of node pairs output by the model: Among them, y uv Indicates whether the node pair (u,v) is connected, is the link probability predicted by the model; E represents the set of real edges in the network.

Citation Information

Cited By

  • Stock equity structure chart identification method and system based on data processing

    CN121031995A

  • A method and system for identifying equity structure diagrams based on data processing

    CN121031995B