An Early Information Propagation Prediction Method and System Based on Diffusion Model
Through the cascading representation learning and time interpolation plus noise mechanism based on diffusion model, the accuracy and real-time prediction of early prediction of information propagation are solved, efficient popularity prediction in the case of scarcity of data is achieved, and the accuracy and computing efficiency of information propagation prediction are improved.
Patent Information
- Application Number
- CN202510125908.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-01-27
AI Technical Summary
In the prediction of the popularity of information dissemination, especially in the early stages and data scarcity, it is difficult to accurately capture the temporal dynamics and structural characteristics of information dissemination, and rely on long-term observation windows and a large number of computing resources, resulting in poor prediction results.
Using a diffusion model-based method, a cascading representation learning module and a time interpolation noise addition module is constructed to build a cascading generation model, combining a time interpolation and an inverse denoising mechanism to generate high-quality cascading representations, reducing dependence on historical data, and improving the accuracy and robustness of early predictions.
In the case of insufficient early data, the temporal dynamics and structural characteristics of information dissemination can be accurately captured, the prediction effect and the real-time nature of the model can be improved, the computing resource consumption can be reduced, and the rapid changes in information dissemination can be adapted to.
Smart Images

Figure CN119558490B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information dissemination technology, and in particular, to an early information dissemination prediction method and system based on a diffusion model. Background Art
[0002] The prediction of the popularity of information dissemination is an important research direction in the current field of social network analysis, and it has wide application value in aspects such as recommendation systems, advertising placement, and social media trend analysis. The process of information dissemination is usually described as an information cascade, which refers to the process of information dynamically spreading through a user network. The goal of popularity prediction is to quantify the attention that a piece of information can obtain, such as the number of forwards or shares. Accurately predicting the popularity of information dissemination can provide a scientific basis for optimizing dissemination strategies, enhancing user engagement, and formulating marketing plans.
[0003] Existing popularity prediction methods are mainly divided into two categories: feature-based methods and deep learning-based methods. Feature-based methods predict by manually extracting key features such as dissemination content, publisher attributes, and publication time. Although such methods can provide a certain prediction ability in the early stage, due to the complexity and insufficiency of feature design, they are often limited in terms of accuracy. In recent years, deep learning methods have made significant progress in information dissemination modeling. In particular, the application of recurrent neural networks and graph neural networks provides powerful tools for capturing the time series characteristics and structural characteristics of dissemination. However, these methods rely on a long observation window, and their prediction ability usually decreases in the early dissemination stage and is difficult to meet the real-time requirements. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide an early information dissemination prediction method based on a diffusion model to eliminate or improve one or more defects existing in the prior art.
[0005] One aspect of the present invention provides an early information dissemination prediction method based on a diffusion model. The steps of the method include:
[0006] Obtain the statistically obtained user social graph and the users currently participating in the information dissemination to be predicted and the time points at which the users participate in the dissemination;
[0007] Determine the user codes of the users participating in the information dissemination to be predicted based on the user social graph, and determine a user sequence group based on the user codes and the time points at which the users participate in the dissemination;
[0008] Sequentially construct a historical dissemination sequence based on each user participating in the information dissemination to be predicted and the time point at which the user participates in the dissemination;
[0009] Input the historical propagation sequence into a pre-trained cascaded representation learning module, and the cascaded representation learning module outputs a first cascaded representation sequence group;
[0010] Input the first cascaded representation sequence group into a pre-trained cascaded generation model, and the cascaded generation model outputs a second cascaded representation sequence group. The cascaded generation model has the structure of a diffusion model;
[0011] Input the second cascaded representation sequence group into a pre-trained popularity prediction module to output a predicted popularity value.
[0012] Adopting the above solution, traditional deep learning methods rely on a large amount of historical data and long observation windows, resulting in unsatisfactory prediction effects in the initial stage of information dissemination. This solution constructs a cascaded generation model based on the diffusion process, adopts a time interpolator and a reverse denoising mechanism, and can generate high-quality cascaded representations. In the case of insufficient early data, this solution can still capture the temporal dynamics and structural features of the cascade, thereby providing more accurate input features for popularity prediction and significantly improving the prediction effect and model robustness in the initial stage of information dissemination.
[0013] In some embodiments of the present invention, the user social graph is a user relationship represented by an adjacency matrix. The adjacency matrix includes rows and columns corresponding to each user. Based on the user social graph, determine the user encoding of the users participating in the dissemination of the information to be predicted, and use the rows of the users participating in the dissemination of the information to be predicted in the adjacency matrix as the user encoding.
[0014] In some embodiments of the present invention, in the step of determining the user sequence group based on the user encoding and the time points at which the user participates in the dissemination, encode the time points to obtain time stamp encodings, and splice the user encoding and the time stamp encodings to obtain the user sequence group.
[0015] In some embodiments of the present invention, the structure of the cascaded representation learning module is a series architecture of a Transformer and a graph attention network.
[0016] In some embodiments of the present invention, the cascaded generation model includes a sequentially connected time interpolation and noise addition module and a reverse denoising module. The time interpolation and noise addition module includes multiple noise addition sub-modules, and the reverse denoising module includes multiple denoising sub-modules. The time interpolation and noise addition module outputs a user sequence group at the target prediction time point, and the reverse denoising module is used to update the user sequence group at the target prediction time point and use the finally updated user sequence group at the target prediction time point as the second cascaded representation sequence group.
[0017] In some embodiments of the present invention, multiple noise addition sub-modules in the time interpolation and noise addition module are connected in series in sequence, and each noise addition sub-module is used to predict the user sequence group at the next time point.
[0018] In some embodiments of the present invention, multiple denoising sub-modules in the reverse denoising module are connected in series in sequence. Each denoising sub-module includes an interpolation network and a prediction network. The interpolation network is used to interpolate the user sequence group between the first concatenated representation sequence group and the user sequence group at the target prediction time point to obtain an interpolated user sequence group, and input the interpolated user sequence group into the prediction network. The prediction network updates the user sequence group at the target prediction time point based on the interpolated user sequence group.
[0019] In some embodiments of the present invention, the network structures of the interpolation network and the noise addition sub-module are the same, both being a convolutional layer, an attention layer, an embedding layer, an attention layer, an embedding layer, an attention layer, and a convolutional layer connected in sequence.
[0020] In some embodiments of the present invention, in the step of inputting the second concatenated representation sequence group into the pre-trained popularity prediction module and outputting the predicted popularity value, the popularity prediction module uses an MLP network.
[0021] The second aspect of the present invention also provides an early information dissemination prediction system based on a diffusion model. The system includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method described above.
[0022] The third aspect of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps implemented by the aforementioned early information dissemination prediction method based on a diffusion model.
[0023] The additional advantages, objectives, and features of the present invention will be partially elaborated in the following description, and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be pointed out and obtained specifically in the description and the accompanying drawings.
[0024] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above specifically described, and it will be more clearly understood from the following detailed description that the above and other objectives that the present invention can achieve. Description of the Drawings
[0025] The accompanying drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0026] Figure 1 It is a schematic diagram of an embodiment of the method for predicting early information dissemination based on the diffusion model of the present invention;
[0027] Figure 2 It is a schematic diagram of a processing architecture of the method for predicting early information dissemination based on the diffusion model of the present invention;
[0028] Figure 3 It is another schematic diagram of a processing architecture of the method for predicting early information dissemination based on the diffusion model of the present invention;
[0029] Figure 4 It is a schematic diagram for comparing the present invention with the prior art;
[0030] Figure 5 It is a performance curve of the prior art and the present solution in the observation time sensitivity analysis;
[0031] Figure 6 It is a curve of the prior art and the present solution in the hyperparameter sensitivity analysis;
[0032] Figure 7 It is a schematic diagram of the ablation results of the prior art and the present solution in the prediction task. Detailed implementation manners
[0033] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with the implementation manners and the accompanying drawings. Herein, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0034] Herein, it also needs to be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the accompanying drawings, while other details less related to the present invention are omitted.
[0035] Introduction to the prior art:
[0036] In the field of information dissemination popularity prediction, the prior art mainly can be divided into feature-based methods and deep learning-based methods.
[0037] Feature-based methods predict by extracting key features in information dissemination (such as the quality of content, the influence of publishers, dissemination time, etc.) and combining traditional machine learning algorithms. Such methods can quickly give popularity predictions in the initial stage of information dissemination. Because they do not rely on a large amount of historical data support, they have high real-time performance. For example, a feature-based popularity prediction method can predict popularity in the initial stage of information dissemination by analyzing the content of information and user characteristics, and is suitable for scenarios with less data volume. Nevertheless, these methods rely heavily on feature selection, require manual extraction and optimization of features, and due to their relatively simple modeling, they cannot fully capture complex dissemination patterns, so the accuracy is relatively low.
[0038] Deep learning-based methods can capture the time dependence in the information dissemination process and the structural relationship between users by automatically learning the latent patterns in time series data and graph structure data, thus greatly improving the prediction accuracy. For example, the DeepHawkes model uses GRU to model the time dynamics in the information dissemination process, and predicts the dissemination trend of information by effectively capturing the dissemination path; and the CasFlow model, which combines the GNN and GRU network structures, can consider both the structural characteristics and time evolution of the dissemination graph, significantly improving the accuracy of popularity prediction.
[0039] Disadvantages of the prior art
[0040] As Figure 4 shown, although the existing feature-based methods can make real-time predictions in the case of scarce data, they face the problems of difficult feature selection and low accuracy. Feature selection is crucial for feature-based methods, but this process usually requires manual design and optimization by domain experts and cannot comprehensively capture the complex relationships and deep features in the information dissemination process. For example, although the feature-based popularity prediction method can handle short-term dissemination well, in complex and long-term dissemination processes, the limitations of feature selection result in relatively low prediction accuracy. In addition, feature-based methods lack in-depth modeling of the dissemination process and cannot effectively capture the time dynamics in the information dissemination process and the complex interaction relationships between users. Therefore, in long-term dissemination predictions, the accuracy performance is poor.
[0041] Although deep learning-based methods can automatically extract features and perform complex modeling, they rely on a large amount of historical data for training and require a long observation window to obtain relatively accurate predictions. In the initial stage of information dissemination, due to the lack of sufficient historical data, deep learning-based models often cannot provide efficient real-time predictions. Especially in rapidly changing environments such as social media, although the CasFlow model considers graph structure characteristics and propagation dynamics, it also requires a large amount of historical data for training and cannot provide accurate predictions when the amount of data is small. In addition, the training process of these deep learning models requires a large amount of computing resources and a long training time. Especially when dealing with large-scale social network data, the computational complexity is high, which limits their feasibility in real-time application scenarios.
[0042] Therefore, a new method needs to be proposed that can generate more accurate and reliable prediction results in the case of scarce data and short observation windows.
[0043] As Figure 1 、 2 and shown in Figure 3, the present invention proposes an early information dissemination prediction method based on a diffusion model. The steps of this method include:
[0044] Step S100, obtaining a statistically obtained user social graph and the users currently participating in the information dissemination to be predicted and the time points at which these users participate in the dissemination.
[0045] In the specific implementation process, the user social graph indicates whether there is a social relationship between users. Specifically, an adjacency matrix can be used to represent the social relationship graph. The rows and columns in the adjacency matrix correspond to users. If there is a social relationship between two users, a value of 1 is set at the intersection of the row and column, otherwise it is 0.
[0046] In the specific implementation process, the user social graph is obtained based on pre-statistical data. Specifically, it can be constructed by staff based on processing experience for evaluation; it can also be the number of interactions between users. The number of interactions between each pair of users is constructed as an input vector, and an adjacency matrix is output through a neural network model.
[0047] Step S200, determining the user codes of the users participating in the information dissemination to be predicted based on the user social graph, and determining a user sequence group based on the user codes and the time points at which these users participate in the dissemination.
[0048] In the specific implementation process, in the step of determining the user codes of the users participating in the information dissemination to be predicted based on the user social graph, the values of the row corresponding to the user in the adjacency matrix are constructed as the user codes.
[0049] Step S300: Based on each user participating in the propagation of the information to be predicted and the time point when the user participates in the propagation, construct a historical propagation sequence in sequence;
[0050] In the specific implementation process, the user participating in the propagation of the information to be predicted and the time point when the user participates in the propagation are taken as a user group, and each user group is sorted based on the order of the time points to obtain the historical propagation sequence.
[0051] Adopting the above solution, this solution captures the temporal features and structural features of users in the information propagation process, and then generates a propagation representation of the information cascade.
[0052] Step S400: Input the historical propagation sequence into a pre-trained cascade representation learning module, and the cascade representation learning module outputs a first cascade representation sequence group;
[0053] In the specific implementation process, the first cascade representation sequence group corresponds to the last user group in the historical propagation sequence, and the last user group corresponds to the last time point;
[0054] Specifically, assume a given user set and an observed historical cascade set . Among them, the historical propagation sequence can be expressed as , which records the propagation process of the information item in ascending order of time, indicating that each user forwards the information item at the time point .
[0055] Step S500: Input the first cascade representation sequence group into a pre-trained cascade generation model, and the cascade generation model outputs a second cascade representation sequence group, and the cascade generation model is of the structure of a diffusion model;
[0056] Step S600: Input the second cascade representation sequence group into a pre-trained popularity prediction module, and output the predicted popularity value.
[0057] Adopting the above solution, traditional deep learning methods rely on a large amount of historical data and a long observation window, resulting in unsatisfactory prediction effects in the initial stage of information propagation; this solution constructs a cascade generation model based on the diffusion process, adopts a time interpolator and a reverse denoising mechanism, and can generate high-quality cascade representations. This solution can still capture the time dynamics and structural features of the cascade in the case of insufficient early data, thereby providing more accurate input features for popularity prediction and significantly improving the prediction effect and model robustness in the initial stage of information propagation.
[0058] In addition, the present invention has also achieved a significant improvement in computational efficiency. By introducing an efficient time dynamic interpolation mechanism, the model can quickly generate future cascade representations in a shorter time, reducing the dependence on long observation windows and effectively reducing the consumption of computing resources. At the same time, the generation process of the present invention adopts a strategy of step-by-step prediction and interpolation sampling, greatly improving the real-time adaptation ability of the model, being able to quickly respond to changes in new information propagation cascade data, and providing efficient and accurate popularity prediction.
[0059] In some embodiments of the present invention, the user social graph is a user relationship represented by an adjacency matrix. The adjacency matrix includes rows and columns corresponding to each user. Based on the user social graph, a user encoding of the users participating in the propagation of the information to be predicted is determined, and the rows of the users participating in the propagation of the information to be predicted in the adjacency matrix are used as the user encoding.
[0060] In some embodiments of the present invention, in the step of determining the user sequence group based on the user encoding and the time points at which the user participates in the propagation, the time points are encoded to obtain timestamp encodings, and the user encoding and the timestamp encodings are concatenated to obtain the user sequence group.
[0061] In some embodiments of the present invention, the structure of the cascade representation learning module is a tandem architecture of a Transformer and a graph attention network.
[0062] In the specific implementation process, an architecture based on a Transformer and a graph attention network (GAT) is adopted, which are respectively used to model the context dependence of users in the cascade and the structural information in the social graph;
[0063] In some embodiments of the present invention, the cascade representation learning module aims to learn the representation of the input cascade by utilizing the user interactions in the social graph and the diffusion cascade, so as to capture the temporal characteristics and the interaction relationship between users in the information propagation process. To achieve this goal, an architecture based on a Transformer and a graph attention network (GAT) is adopted, which are respectively used to model the context dependence of users in the cascade and the structural information in the social graph.
[0064] In some embodiments of the present invention, the cascade generation model includes a sequentially connected time interpolation and noise addition module and a reverse denoising module. The time interpolation and noise addition module includes multiple noise addition sub-modules, and the reverse denoising module includes multiple denoising sub-modules. The time interpolation and noise addition module outputs the user sequence group at the target prediction time point, and the reverse denoising module is used to update the user sequence group at the target prediction time point, and the finally updated user sequence group at the target prediction time point is used as the second cascade representation sequence group.
[0065] In some embodiments of the present invention, multiple noise addition sub-modules in the time interpolation and noise addition module are connected in series in sequence, and each noise addition sub-module is used to predict the user sequence group at the next time point.
[0066] With the above solution, this solution focuses on the macroscopic level of information dissemination through the time interpolation and noise addition module. By generating cascaded future states, it estimates the overall scale of information dissemination; this module simulates the noise addition process during the diffusion process. First, it smooths the time dynamics in historical data through time-conditioned interpolation to reduce discontinuities during the dissemination process; second, it introduces appropriate randomness through the noise addition process to simulate the uncertainty and complexity in information dissemination, so as to enhance the robustness of the model and its adaptability to the dissemination trend.
[0067] In some embodiments of the present invention, multiple denoising sub-modules in the reverse denoising module are connected in series in sequence. Each denoising sub-module includes an interpolation network and a prediction network. The interpolation network is used to interpolate the user sequence group between the first cascaded representation sequence group and the user sequence group at the target prediction time point to obtain the interpolated user sequence group, and input the interpolated user sequence group into the prediction network. The prediction network updates the user sequence group at the target prediction time point based on the interpolated user sequence group.
[0068] In the specific implementation process, in the diffusion process modeling, the present invention uses the Transformer structure to model the interactions of users in each cascade.
[0069] In the specific implementation process, at each time step, the interpolation network is used to smooth the current prediction result and generate an intermediate representation to ensure the continuity between time steps. The specific process is executed iteratively to gradually advance the time evolution until the cascaded representation at the target time step is generated.
[0070] With the above solution, the interpolation network introduces appropriate randomness to simulate the uncertainty and complexity in information dissemination, so as to enhance the robustness of the model and its adaptability to the dissemination trend; the prediction network simulates the denoising process during the diffusion process. Based on the future dissemination state generated by time interpolation, a popularity increment predictor for information dissemination is trained.
[0071] In the specific implementation process, this solution also includes pre-training the cascade generation model and using the mean square error as the loss function to train the cascade generation model.
[0072] In some embodiments of the present invention, the network structures of the interpolation network and the noise addition sub-module are the same, both being sequentially connected convolutional layers, attention layers, embedding layers, attention layers, embedding layers, attention layers, and convolutional layers.
[0073] Specifically, the input information is gradually encoded through a convolutional layer, a temporal embedding layer, and a residual cross-attention mechanism.
[0074] In some embodiments of the present invention, in the step of inputting the second concatenated representation sequence group into a pre-trained popularity prediction module to output a predicted popularity value, the popularity prediction module employs an MLP network.
[0075] In the specific implementation process, a multi-layer perceptron (MLP) is used as the prediction network. By learning the relationship between the concatenated representation and the popularity increment, the prediction result is finally obtained.
[0076] In the specific implementation process, this solution further includes pre-training the popularity prediction module, using the mean squared logarithmic error (MSLE) as the loss function. Specifically, by minimizing this loss function, the popularity prediction module can accurately learn the mapping relationship between the cascading evolution and the final popularity, and achieve the prediction of future popularity.
[0077] In summary, the present invention solves the limitations of the prior art in the popularity prediction task of information dissemination when facing problems such as data scarcity, strong time dynamics, and high real-time requirements. Although the existing deep learning-based methods can better model complex time series and graph-structured data, their dependence on a large amount of historical data and long observation windows makes these methods have poor prediction effects in data-scarce scenarios, especially at the beginning of information dissemination. Therefore, the present invention proposes a novel generative information dissemination prediction method, which overcomes the challenges of data insufficiency and early prediction by generating concatenated representations, thereby providing more accurate popularity predictions.
[0078] The deep learning methods of the prior art usually require a long training time and a large amount of computing resources for popularity prediction, which makes them unable to exert the best performance in scenarios with high real-time requirements. In the initial stage of information dissemination, since the disseminated information has not been fully unfolded, traditional models are difficult to quickly adapt and provide effective predictions. Therefore, the present invention designs an efficient generative model, reduces the dependence on a large amount of historical data, and at the same time, by introducing time dynamic interpolation based on the diffusion process, enables the model to quickly adapt to new dissemination cascade data and improves the accuracy and efficiency of real-time prediction.
[0079] Experimental Example:
[0080] I. Experimental Settings
[0081] 1. Dataset
[0082] This experimental example evaluates the proposed framework on three public datasets: Twitter, Weibo, and APS.
[0083] In this experimental example, 70% of the cascades were randomly selected as the training set, 15% as the validation set, and the remaining 10% as the test set. The detailed information of the dataset is shown in Table 1:
[0084] Table 1
[0085]
[0086] 2. Comparison methods
[0087] In this experimental example, this solution was compared with the following existing technologies, which include:
[0088] DeepHawkes: Each cascade is modeled as a set of diffusion paths between users, and a gated recurrent unit (GRU) is used to capture the sequential progress of the cascade.
[0089] MS-HGAT: A series of time-sampled hypergraphs are constructed to encapsulate multiple cascades and users, and hypergraph learning is used to calculate the cascade representation.
[0090] CasCN: Each cascade is modeled as a time graph sequence, and a combination of a graph neural network (GNN) and a long short-term memory (LSTM) network is used to learn a robust cascade representation.
[0091] TempCas: Integrate specialized sequence modeling techniques, aiming to capture the overall temporal patterns and supplement its learning on the cascade graph.
[0092] CasFlow: First, user representations are derived from the social network and the cascade graph, and then GRU combined with a variational autoencoder (VAE) is used to encode the cascade representation.
[0093] CTCP: The state-of-the-art cascade prediction method that groups multiple cascades by sharing propagation users and enables a unified temporal and structural learning process across cascades.
[0094] 3. Evaluation metrics
[0095] In this experimental example, the following four widely recognized metrics were used to evaluate the performance of the comparison methods:
[0096] Mean Squared Logarithmic Error (MSLE): Quantify the prediction error by calculating the mean of the squared differences between the logarithms of the predicted values and the true values. A lower MSLE value indicates a smaller logarithmic difference between the model prediction and the true value, and higher prediction accuracy.
[0097] Mean Absolute Logarithmic Error (MALE): The average of the absolute differences between the logarithms of the predicted values and the true values. A lower MALE value means a smaller average difference between the predicted value and the true value on the logarithmic scale.
[0098] Mean Absolute Percentage Error (MAPE): It measures the percentage of prediction error by calculating the average of the ratios of the absolute values of the differences between the predicted values and the true values to the true values. A lower MAPE value indicates a smaller average percentage difference between the predicted values and the true values, and a higher prediction accuracy.
[0099] Pearson Correlation Coefficient (PCC): It measures the linear correlation between the predicted values and the true values. A higher PCC value (close to 1) indicates a strong positive correlation between the predicted values and the true values, indicating that the model can better capture the trends and patterns in the data.
[0100] II. Performance Evaluation
[0101] 1. Model Performance
[0102] To evaluate the performance of this solution under different datasets, performance analysis was conducted on Twitter, Weibo, APS data, etc. The results are shown in Table II:
[0103] Table II
[0104]
[0105] The following important conclusions can be drawn from the results in Table II:
[0106] Feature-based models, such as DeepHawkes, lag behind other methods in all evaluation metrics. This can be attributed to the inherent limitations of feature-based models in capturing the complex non-linear evolution patterns of cascade sizes, especially in the early stages of information propagation.
[0107] Graph-based models are generally superior to sequence-based models, which precisely illustrates the importance of integrating structural and temporal information in cascade graphs. For example, CasFlow and CTCP show good performance on the Weibo dataset, indicating their effectiveness in modeling the temporal and structural dynamics of information propagation. However, their performance is not outstanding within a short observation period.
[0108] This solution outperforms other models in all metrics, especially on the Twitter and Weibo datasets with a short observation window, which highlights its ability to generate accurate predictions with less initial data.
[0109] 2. Observation Time Sensitivity Analysis
[0110] To further evaluate the performance of this solution under different observation time windows, this experimental example conducted observation time sensitivity analysis on the Twitter and APS datasets and plotted the performance curves of the Top3 models in these intervals.
[0111] AsFigure 5 As shown in the figure, as the observation time window is extended, the PCC values of all models increase. This indicates that a longer observation time window helps the model better capture the underlying dynamics of information dissemination. However, this solution maintains a high PCC value within all observation time windows. Especially in the early stage of information cascades (i.e., within a shorter observation time window), this solution significantly outperforms other models, which shows that this solution can effectively capture the initial dynamics of information dissemination and can better model cascade dynamics in the case of sparse data.
[0112] 3. Hyperparameter Sensitivity Analysis
[0113] To evaluate the sensitivity of the model to hyperparameters, this experimental example conducts a sensitivity analysis on two key hyperparameters: the time step (Timestep) and the number of iterations (Horizon) during the sampling process.
[0114] From Figure 6 it can be seen that if the time step is too large, the generated embeddings may not be consistent with the true cascade evolution distribution. Additionally, too many iterations may introduce noise, thus reducing the effectiveness of the model.
[0115] III. Ablation Experiments
[0116] To evaluate the importance of each module in this solution, this experimental example conducts ablation experiments on the Twitter and APS datasets. This experimental example removes the time learning module (TL), the structural representation module (SL), and the cascade generation module (GM) respectively, and evaluates the performance of these variants. The relevant results are as Figure 7 shown.
[0117] Specifically, removing the time learning module means that there are no time points in the user sequence group that are propagated through user participation; the obtained encoding of time;
[0118] Removing the structural representation module means that there are no user encodings obtained from the user social graph in the user sequence group;
[0119] Removing the cascade generation module (GM) means removing the cascade generation model.
[0120] As Figure 7 shown, removing the time learning module (TL): Removing the TL module will cause a significant decrease in MSLE and PCC, especially on the Twitter dataset. This indicates that time dynamics play a key role in the early stage of information dissemination, where the timing and rhythm of forwarding or quoting are key indicators of the popularity of future cascades. The time features captured by this module are particularly important in the case of a short observation period, such as in a social media environment.
[0121] Removal of the Structure Representation Module (SL): Removing the SL module also leads to a performance degradation, but its impact is relatively smaller than that of the TL module, especially on the APS dataset. In a network like APS, the structure changes slowly, and the temporal evolution may play a more important role in predicting future cascade growth.
[0122] Removal of the Cascade Generation Module (GM): Removing the GM module results in a significant performance degradation, indicating that the cascade generation module is crucial in this application. Its ability to simulate future cascade propagation based on limited initial data is essential for long-term prediction. Its significant contribution to the MSLE and PCC metrics highlights the importance of generating synthetic data.
[0123] In summary, this solution proposes a novel generation method for the information propagation prediction task, aiming to address the challenges of early information popularity prediction. This solution draws on the gradual noise addition and denoising processes of diffusion models, using a time-conditioned interpolator as the noise addition process and a denoising process for prediction, iteratively generating cascade representations that can effectively capture the rich temporal dynamics of information propagation. Subsequently, in the sampling stage, cascade representations are generated, which are input into the prediction module to generate popularity predictions. Experimental examples demonstrate that the comprehensive evaluation of this solution on different datasets such as Twitter, Weibo, and APS highlights the effectiveness of this solution, indicating that its performance in early prediction tasks is superior to state-of-the-art methods. This work not only improves the accuracy and timeliness of information popularity prediction but also has important potential value for applications such as recommendation systems, targeted advertising, and social media trend analysis.
[0124] The beneficial effects of the present invention include:
[0125] 1. By combining a time interpolator and a reverse denoiser, the present invention models the information propagation process as a diffusion generation process. In the forward diffusion, the smooth generation of cascade representations at different time steps is achieved through the time interpolator, thereby capturing temporal dynamics and early propagation characteristics. In the reverse denoising process, future cascade representations are generated by gradually denoising, effectively solving the problems of data scarcity and dependence on long observation windows;
[0126] 2. The present invention proposes a generation and sampling mechanism that combines interpolation and denoising. Through step-by-step prediction and interpolation correction, the model can generate future cascade representations in a short time, reducing the dependence on long historical data. At the same time, this mechanism ensures the temporal smoothness and continuity of the generation process through continuous modeling of time steps, thereby improving the prediction efficiency and real-time performance of the model.
[0127] 3. The present invention does not rely on a large amount of historical data and can achieve relatively accurate cascaded feature learning in the case of sparse data or a small observation window, and can meet the timeliness requirements of computational efficiency.
[0128] An embodiment of the present invention further provides an early information dissemination prediction system based on a diffusion model. The system includes a computer device, the computer device includes a processor and a memory, computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps implemented by the method described above.
[0129] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps implemented by the foregoing early information dissemination prediction method based on a diffusion model. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium well-known in the technical field.
[0130] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0131] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0132] In the present invention, features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0133] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An early information dissemination prediction method based on a diffusion model, characterized in that, The steps of the method include: Obtaining the statistically obtained user social graph, the users currently participating in the dissemination of the information to be predicted, and the time points at which these users participate in the dissemination; Determining the user codes of the users participating in the dissemination of the information to be predicted based on the user social graph, and determining a user sequence group based on the user codes and the time points at which these users participate in the dissemination; Sequentially constructing a historical dissemination sequence based on each user participating in the dissemination of the information to be predicted and the time point at which this user participates in the dissemination; Inputting the historical dissemination sequence into a pre-trained cascaded representation learning module, and the cascaded representation learning module outputs a first cascaded representation sequence group; Inputting the first cascaded representation sequence group into a pre-trained cascaded generation model, and the cascaded generation model outputs a second cascaded representation sequence group. The cascaded generation model has the structure of a diffusion model. The cascaded generation model includes a sequentially connected time interpolation noise addition module and a reverse denoising module. The time interpolation noise addition module includes multiple noise addition sub-modules, and the reverse denoising module includes multiple denoising sub-modules. The time interpolation noise addition module outputs a user sequence group at the target prediction time point. The reverse denoising module is used to update the user sequence group at the target prediction time point. The multiple denoising sub-modules in the reverse denoising module are sequentially connected in series. Each denoising sub-module includes an interpolation network and a prediction network. The interpolation network is used to interpolate the user sequence group between the first cascaded representation sequence group and the user sequence group at the target prediction time point. The prediction network updates the user sequence group at the target prediction time point based on the interpolated user sequence group, and takes the finally updated user sequence group at the target prediction time point as the second cascaded representation sequence group; Inputting the second cascaded representation sequence group into a pre-trained popularity prediction module, and outputting a predicted popularity value.
2. The early information dissemination prediction method based on the diffusion model according to claim 1, wherein The user social graph is a user relationship represented by an adjacency matrix. The adjacency matrix includes rows and columns corresponding to each user. Determining the user codes of the users participating in the dissemination of the information to be predicted based on the user social graph, and taking the rows of the users participating in the dissemination of the information to be predicted in the adjacency matrix as the user codes.
3. The early information dissemination prediction method based on a diffusion model according to claim 1, characterized in that In the step of determining a user sequence group based on the user codes and the time points at which these users participate in the dissemination, encoding the time points to obtain time stamp encodings, and concatenating the user codes and the time stamp encodings to obtain a user sequence group.
4. The early information dissemination prediction method based on the diffusion model according to claim 1, wherein The structure of the cascaded representation learning module is a cascaded architecture of a Transformer and a graph attention network.
5. The early information dissemination prediction method based on a diffusion model according to claim 1, wherein The multiple noise addition sub-modules in the time interpolation noise addition module are sequentially connected in series, and each noise addition sub-module is used to predict the user sequence group at the next time point.
6. The early information dissemination prediction method based on a diffusion model according to claim 1, characterized in that The network structures of the interpolation network and the noise addition sub-module are the same, and both are sequentially connected convolutional layers, attention layers, embedding layers, attention layers, embedding layers, attention layers, and convolutional layers.
7. The early information dissemination prediction method based on the diffusion model according to claim 1, wherein In the step of inputting the second cascaded representation sequence group into a pre-trained popularity prediction module and outputting a predicted popularity value, the popularity prediction module uses an MLP network.
8. An early information dissemination prediction system based on a diffusion model, characterized in that, The system includes a computer device, the computer device includes a processor and a memory, computer instructions are stored in the memory, the processor is configured to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the system implements the steps implemented by the method according to any one of claims 1 to 7.