An Information Cascade Prediction Method Based on Multi-Task Learning
Through the multi-task learning method, combining long and short-term memory networks and bidirectional graph convolution neural networks, and fusing propagation characteristics at the micro and macro levels, the problem of lack of correlation between information cascade prediction methods in the existing technology is solved, and more efficient prediction performance and interpretability are achieved.
Patent Information
- Application Number
- CN202210817949.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-07-12
AI Technical Summary
The existing information cascade prediction methods fail to effectively combine the correlation between the micro and macro levels, resulting in insufficient performance when predicting information dissemination trends and lack of interpretability and migration.
The multi-task learning method is adopted to model the propagation characteristics at the micro and macro levels through long and short-term memory networks and improved bidirectional graph convolutional neural networks, and use a shared gating mechanism to blend the information representation between the two to achieve end-to-end prediction.
It improves the interpretability and migration of information cascading prediction, improves the performance of micro and macro prediction, and can be effectively applied in different fields.
Smart Images

Figure CN115392431B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information cascade prediction, and specifically to an information cascade prediction method based on multi-task learning. Background Art
[0002] The information dissemination pattern in an online network is complex and changeable. Different dissemination processes can generate various complex network models. Modeling the network model and predicting the future dissemination trend of information have become research issues that have attracted much attention. When a piece of information is created and published by an original user at a certain moment, other users can view the information through various relationship chains and perform operations such as liking, commenting, or forwarding it, thereby participating in the dissemination process of the information. The forwarding by forwarding users causes the information to be forwarded by more users through the same dissemination process, resulting in a hierarchical progressive forwarding of the original information. The overall forwarding process will constitute the information cascade corresponding to the original information.
[0003] For an information cascade, existing methods model the information dissemination process and capture the dissemination pattern to predict the future dissemination trend of information. According to different research perspectives, the information cascade prediction task can be divided into a micro-cascade prediction task and a macro-cascade prediction task. The micro-cascade prediction task starts from the user perspective, regards the forwarding process of the information cascade as an independent event, and uses the observable cascade sequence to construct a high-performance dissemination model to predict the users who will participate in the dissemination process at the next moment. The macro-cascade prediction task starts from the perspective of the overall cascade, inputs the cascade within a given observation time window into the macro-cascade prediction model, and predicts the scale that the information cascade can reach in the future.
[0004] Micro-cascade prediction methods can be classified into methods based on prior dissemination models, methods based on representation learning, and methods based on neural networks. Methods based on prior dissemination models rely on pre-set information dissemination probabilities and patterns, can describe the information dissemination process, but their effectiveness depends to a large extent on the assumptions of the potential dissemination model, which are difficult to specify or verify in practice and have a large parameter scale. Methods based on representation learning embed users into a potential vector representation space and predict the dissemination probability by calculating the similarity between user vectors. Such methods reduce the parameter scale but completely ignore the infection order of the dissemination sequence and cannot capture the potential characteristics and influences of the information cascade. Methods based on neural networks use models such as recurrent neural networks to automatically learn the historical dissemination path for micro-cascade prediction, and the prediction performance is good, but they focus on the local pattern of information dissemination and cannot capture the global information of the cascade.
[0005] Macro - level cascade prediction methods can be divided into feature - extraction - based methods, generative methods, and deep - learning - based methods. Feature - extraction - based methods use prior knowledge to extract and construct features from the original data as much as possible, and use machine - learning models to establish a mapping between the features and the popularity labels to be predicted ultimately. Such methods reveal various effective factors affecting the evolution of popularity. However, the extraction of features requires strong domain experience, and the manually extracted features may only be applicable to specific datasets or specific scenarios, with poor generality. At the same time, feature - extraction - based methods only fit the final popularity by constructing features, and cannot reveal complex propagation dynamics, with poor interpretability. Generative methods model the intensity function of each piece of information during its occurrence process, can depict the dynamic growth process of information, and have strong interpretability. However, due to not being supervised by future popularity during the training process, the prediction performance of such methods is low. Deep - learning - based methods aim to model the propagation cascade graph, use the deep - learning framework to automatically learn valuable representations from the data, and map them to the final popularity for prediction. Such methods achieve better prediction performance, but regard the propagation graph as a static graph and cannot capture the dynamic evolution process of the entire propagation graph.
[0006] Micro - level cascade prediction models and macro - level cascade prediction models have both achieved certain results at different levels of cascade modeling. However, most of the existing cascade prediction methods ignore the mutual relevance between the two types of tasks, treat the micro - prediction task and the macro - prediction task as independent tasks and process them separately, and cannot capture the rich correlation information contained between the two types of tasks. Currently, there is a lack of a method in the field of information cascade prediction that simultaneously models the propagation process by combining the micro and macro levels. Summary of the Invention
[0007] Aiming at the defects existing in the prior art, the purpose of the present invention is to provide an information cascade prediction method based on multi - task learning. The method provided by the present invention integrates the idea of multi - task learning, that is, uses multi - task learning to capture the propagation characteristics of the cascade from both the micro and macro levels simultaneously, and models the interaction between them. This method has good interpretability, verifiability, and transferability, and has important theoretical research value and application value.
[0008] To achieve the above - mentioned purpose, the technical solution adopted by the present invention is:
[0009] An information cascade prediction method based on multi - task learning, characterized by comprising the following steps:
[0010] S1. Pre - process the original information cascade, arrange the users participating in the propagation according to the forwarding time to obtain a forwarding sequence, and slice the cascade graph according to the time interval to obtain a snapshot sequence composed of snapshots at each slice moment;
[0011] S2. Initialize the feature representation of each user in the forwarding sequence obtained in step S1 using one-hot encoding, and input it into a long short-term memory network to obtain the corresponding microscopic information representation H user ;
[0012] S3. Model the snapshot sequence obtained in step S1 using the method of a temporally improved bidirectional graph convolutional neural network to obtain a macroscopic information representation H containing cascaded dynamic evolution information cas ;
[0013] S4. Use an improved shared gating mechanism to fuse the above microscopic information representation H user and macroscopic information representation H cas to obtain a shared representation H that fuses information from two levels share ;
[0014] S5. Concatenate the shared representation H obtained in step S4 with the information representations corresponding to the tasks respectively. Input the microscopic concatenated representation in the above information representation concatenation into the softmax function to predict the next user participating in the propagation, and input the macroscopic concatenated representation in the above information representation concatenation into a multi-layer perceptron to predict the popularity increment Δy share On the basis of the above scheme, the specific steps of step S1 are as follows:
[0015] Given the original information cascade G, arrange the users participating in the propagation according to the forwarding time to obtain a user forwarding sequence S = {v1, v2,..., v
[0016] }, where v N represents each user, and N represents the number of users. Divide the given observation window T into n time intervals, and count the propagation graph information at each slice moment to obtain a snapshot g * = {u i , e i , t i}, where u i , e i , and t i represent the set of users, the set of edges, and the set of forwarding timestamps in the i-th snapshot respectively. Finally, obtain a snapshot sequence G = {g1, g2,..., g i}. n}
[0017] On the basis of the above scheme, the specific steps of step S2 are as follows:
[0018] Given the forwarding sequence S = {v1, v2,..., v N}, initialize the feature representation of each user using one-hot encoding. All users will be associated with a specific embedding matrix E ∈ R N×DPerform association, where D represents the embedding dimension of the vector. For user u, use the embedding matrix E to map it into the corresponding low-dimensional dense representation vector x = uE. Use the long short-term memory network LSTM to model the forwarding sequence after vector representation. Assume that the embedding vector of the user at time t is x i , and the hidden layer state at time t-1 is h i-1 . The calculation method of the hidden layer state at time t is shown in Equation (1):
[0019] h i = LSTM(x i , h i-1 ) (1);
[0020] where d seq represents the dimension of the hidden layer unit.
[0021] Using the update rule of LSTM, the micro-level user sequence H user can be composed of all hidden layer states.
[0022] On the basis of the above scheme, the specific steps of step S3 are as follows:
[0023] Given the snapshot sequence G = {g1, g2,..., g n}, for the i-th snapshot g i , use the TGAT method to encode the timestamp set t i in the snapshot to obtain the timing matrix C i . On the basis of obtaining the adjacency matrix A of the cascaded graph corresponding to the i-th snapshot g i , respectively input the degree matrix D i in g i and the previous timing matrix C i into the bidirectional graph convolutional neural network BiGCN to obtain the corresponding structural information representation and timing information representation After splicing the two types of information representations, the representation corresponding to each snapshot is obtained . The calculation method is shown in Equation (2):
[0024]
[0025]
[0026]
[0027] where I N is the identity matrix, and W C and W D represent trainable parameter matrices.
[0028] The representation of each snapshot can form a snapshot representation sequence Input it into a gated neural network to extract the temporal dependencies between different snapshots, and obtain the representation matrix H of the final cascaded graph cas .
[0029] Based on the above scheme, the specific steps of step S4 are as follows:
[0030] Given the microscopic information representation H user and the macroscopic information representation H cas , use the designed shared gating mechanism to aggregate the information at both levels and obtain the shared representation H containing all the information share . The calculation process of the proposed shared gating mechanism is as follows: First, the forget gate f t is used to select the irrelevant information in the previous state that needs to be forgotten and update the cell state c t , and the calculation method is shown in Equation (3):
[0031]
[0032] where, W f , U f and b f represent the learnable parameters during the training of the gated neural network, and represent the microscopic information representation and the macroscopic information representation at time t - 1 respectively;
[0033] Use the obtained forget gate to filter the information that needs to be forgotten in the cell state at time t - 1, and update to obtain the new cell state c t , and the calculation method is shown in Equation (4), W c and U c represent the learnable parameters during the training of the gated neural network;
[0034]
[0035] The reset gate r t is used to determine the information in the newly input hidden layer state h t , and the calculation method is shown in Equation (5), W r , U r and b r represent the learnable parameters during the training of the gated neural network;
[0036]
[0037] Use the obtained reset gate to perform information selection on the new cell state c t and update to obtain the new hidden layer state ht , the calculation method is as shown in Equation (6), where W h and U h represent the learnable parameters during the training process of the gated neural network;
[0038]
[0039] Utilize the above shared gating mechanism to fuse H user and H cas , and use the hidden state of the last step of the gated recurrent neural network as the shared representation layer H share ;
[0040] In the above Equations (3)-(6), σ(·) represents the sigmoid activation function, and ⊙ represents element-wise multiplication.
[0041] Based on the above solution, the specific steps of step S5 are as follows:
[0042] Given the shared representation layer H share , concatenate the specific representation H user and H cas of each task with it respectively, and input the concatenated specific task representation into different output layers for prediction. The goal of the micro-cascade prediction task is to predict the next user participating in the cascade propagation. The output layer uses the logical classifier softmax function for prediction. For each user, the calculation method of the probability that the model predicts it will participate in the propagation at the i+1 moment is as shown in Equation (7), where concat(·) represents the concatenation operation:
[0043]
[0044] Given the length l of the cascade sequence, the learning goal of the micro-cascade prediction model is to maximize the probability of all users participating in the propagation in the propagation sequence. The objective function is as shown in Equation (8):
[0045]
[0046] The loss function of the micro-prediction task during model training can be regarded as minimizing the cross-entropy loss between the predicted probability of sequence S
[0047]
[0048] and the true probability P, and the calculation method is as shown in Equation (9):
[0047]
[0048] The goal of the macro-cascade prediction task is to predict the popularity increment within a fixed future time interval by observing the cascade information within a fixed time window. The output layer uses the regressor multi-layer perceptron MLP for prediction. For the snapshot sequence G formed by the information cascade, the calculation method of the popularity increment predicted by the model is as shown in Equation (10):
[0049] ΔS = MLP(concat(h cas ,h share )) (10);
[0050] The loss function for the macro prediction task during model training can be regarded as minimizing the error between the predicted value and the true value, and the calculation method is shown in Equation (11):
[0051]
[0052] where Q represents the total number of cascades in the dataset, represents the true popularity increment;
[0053] The overall loss function of the model is shown in Equation (12), where γ ∈ [0, 1] is a learning parameter that balances l1 and l2:
[0054] L = γl1 + (1 - γ)l2 (12).
[0055] An information cascade prediction method based on multi-task learning according to the present invention has the following beneficial effects:
[0056] (1) High interpretability: The present invention uses a shared representation layer based on a gating mechanism to simultaneously capture the underlying information of the propagation cascade graph at the macro level and the propagation node sequence information at the micro level, and learns the mutual influence between the two tasks, having good interpretability.
[0057] (2) Strong transferability: The method in the present invention can achieve cascade prediction at both the micro and macro levels in an end-to-end manner, without involving a large amount of complex feature engineering, and the model is relatively easy to apply to new fields. The described shared gating mechanism can provide support for other multi-task learning methods and has strong transferability.
[0058] (3) Excellent characterization performance: The present invention models the mutual influence between the sequence information at the micro level and the cascade information at the macro level. Through experiments, it is found that the proposed method can simultaneously improve the micro-cascade prediction performance and the macro-cascade prediction performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The present invention has the following drawings:
[0060] Figure 1 is a schematic structural diagram of an information cascade prediction method framework based on multi-task learning;
[0061] Figure 2 is a schematic diagram of a graph convolutional neural network method based on temporal improvement. DETAILED DESCRIPTION OF THE INVENTION
[0062] The present invention will be further described in detail below with reference to the accompanying drawings.
[0063] The present invention proposes an information cascade prediction method based on multi-task learning (hereinafter referred to as the method), as Figure 1 shown. When user A on the Weibo platform posts a Weibo at time t A and then user B forwards this Weibo at time t B and user D is influenced by user B and forwards this Weibo at time t D and the forwarding users are sorted according to the forwarding time to obtain a forwarding sequence (A, B, C, D, E, F). Different users' forwarding of this Weibo has different impacts, resulting in multiple forwarding sequences of different lengths, and the structures of all forwarding sequences will form the cascade graph of this Weibo. Given a fixed-length observation time window (such as 3 hours), at this time the number of times this Weibo is forwarded is 10, and the micro-task predicts the users (G, H, I,...) who will participate in the forwarding of this Weibo at the next moment, and the macro-task predicts the increment of the number of times this Weibo is forwarded between 3 hours and 24 hours.
[0064] First, preprocess the original information cascade to obtain the forwarding sequence for the micro-task and the cascade graph for the macro-task respectively. Second, input the forwarding sequence into the long short-term memory network to obtain the corresponding micro-information representation, and input the cascade graph into the graph convolutional neural network with improved time series (as Figure 2 shown) to obtain the corresponding macro-information representation; third, use the shared gating mechanism to fuse the micro-information representation and the macro-information representation to obtain the shared representation of the information cascade; finally, splice the obtained shared representation with the information representation of the corresponding task and input it into different output layers to realize the prediction of the micro-task and the macro-task, and the specific process is as follows.
[0065] (1) Micro and macro task representations
[0066] In order to more efficiently mine the mutual influence between the micro-task and the macro-task, this method integrates the multi-task learning idea, simultaneously obtains the sequence required for the micro-task and the cascade graph required for the macro-task on the basis of the original data, and performs modeling respectively to obtain the corresponding representations.
[0067] Given the original information cascade G, arrange the users participating in the propagation according to the forwarding time to obtain a user forwarding sequence S = {v1, v2,..., v N}, where v * represents each user and N represents the number of users. Initialize the feature representation of each user using one-hot encoding, and all users will be associated with a specific embedding matrix E ∈ R N×DPerform association, where D represents the embedding dimension of the vector. For user u, use the embedding matrix E to map it into the corresponding low-dimensional dense representation vector x = uE. Use the long short-term memory network LSTM to model the forwarding sequence after vector representation. Assume that the embedding vector of the user at time t is x i , and the hidden layer state at time t - 1 is h i-1 . The calculation method of the hidden layer state at time t is shown in Equation (1). Among them d seq represents the dimension of the hidden layer unit.
[0068] h i = LSTM(x i , h i-1 ) (1)
[0069] Using the update rule of LSTM, the user sequence H user at the micro level can be composed of all hidden layer states.
[0070] Divide the given observation window T into n time intervals, and count the propagation graph information at each slice moment to obtain the snapshot g i = {u i , e i , t i}, where u i , e i and t i represent the set of users, the set of edges, and the set of forwarding timestamps in the i-th snapshot respectively. Finally, obtain the snapshot sequence G = {g1, g2,..., g n}. For the i-th snapshot g i , use the TGAT method to encode the timestamp set t i in the snapshot to obtain the timing matrix C i . On the basis of obtaining the adjacency matrix A of the cascaded graph corresponding to the i-th snapshot g i , respectively input the degree matrix D i in g i and the previous timing matrix C i into the bidirectional graph convolutional neural network BiGCN to obtain the corresponding structural information representation and timing information representation Concatenate the two types of information representations to obtain the representation corresponding to each snapshot The calculation method is shown in Equation (2).
[0071]
[0072]
[0073]
[0074] Among them, I N is the identity matrix, and W C and W D represent trainable parameter matrices.
[0075] The representations of each snapshot can form a snapshot representation sequence Input it into a gated neural network to extract the temporal dependencies between different snapshots, and obtain the representation matrix H of the final cascaded graph cas .
[0076] (2) Shared gating mechanism
[0077] To ensure parameter sharing among multiple tasks, multi-task learning frameworks usually use a shared representation layer to aggregate embedding representations specific to different tasks. To make full use of the representation H user of the user sequence and the representation H cas of the cascaded graph, the method proposes a novel shared gating mechanism, and the specific implementation is as follows:
[0078] Given the microscopic information representation H user and the macroscopic information representation H cas , use the designed shared gating mechanism to aggregate the information at the two levels and obtain the shared representation H share containing all the information. The calculation process of the proposed shared gating mechanism is as follows: First, the forget gate f t is used to select the irrelevant information in the previous state to be forgotten and update the cell state c t , and the calculation method is shown in Equation (3).
[0079]
[0080] Among them, W f , U f and b f represent the learnable parameters in the training process of the gated neural network, and represent the microscopic information representation and the macroscopic information representation at time t - 1 respectively.
[0081] Use the obtained forget gate to filter the information to be forgotten in the cell state at time t - 1 and update to obtain the new cell state c t , and the calculation method is shown in Equation (4). W c and U c represent the learnable parameters in the training process of the gated neural network, σ(·) represents the sigmoid activation function, and ⊙ represents element-wise multiplication.
[0082]
[0083] Reset gate r t Used to determine the input of the new hidden layer state h t The information in is calculated as shown in Equation (5), W r 、U r and b r represent the learnable parameters in the training process of the gated neural network.
[0084]
[0085]
[0086] Utilize the above shared gating mechanism to fuse H user and H cas Take the hidden state of the last step of the gated recurrent neural network as the shared representation layer H of multi-task learning share .
[0087] (3) Specific task learning
[0088] It is to use the shared representation to assist the learning process of different-level tasks respectively, model the mutual influence between different tasks, and a specific task representation method based on multi-task learning is proposed. The specific implementation plan is as follows:
[0089] Given the shared representation layer H share , connect the specific representation H user and H cas of each task with it respectively, and input the concatenated specific task representation into different output layers for prediction. The goal of the micro-cascade prediction task is to predict the next user participating in the cascade propagation, and the output layer uses the logical classifier softmax function for prediction. For each user, the calculation method of the probability that the model predicts that it will participate in the propagation at the i+1 moment is as shown in Equation (7), and concat(·) represents the concatenation operation.
[0090]
[0091] Given the length l of the cascade sequence, the learning goal of the micro-cascade prediction model is to maximize the probability of all users participating in the propagation in the propagation sequence, and the objective function is as shown in Equation (8).
[0092]
[0093] The loss function of the micro-prediction task during model training can be regarded as minimizing the cross-entropy loss between the predicted probability of the sequence S and the true probability P, and the calculation method is as shown in Equation (9).
[0094]
[0095] The goal of the macro-cascade prediction task is to predict the popularity increment within a fixed time interval in the future by observing the cascade information within a fixed time window, and a regressor multi-layer perceptron (MLP) is used in the output layer for prediction. For the snapshot sequence G formed by the information cascade, the calculation method for the model to predict the popularity increment is shown in Equation (10).
[0096] ΔS = MLP(concat(h cas ,h share )) (10)
[0097] The loss function of the macro-prediction task during model training can be regarded as minimizing the error between the predicted value and the true value, and the calculation method is shown in Equation (11).
[0098]
[0099] Among them, Q represents the total number of cascades in the dataset, represents the true popularity increment.
[0100] The loss function of the overall model is shown in Equation (12), where γ ∈ [0, 1] is a learning parameter that balances l1 and l2.
[0101] L = γl1 + (1 - γ)l2 (12)
[0102] The embodiments in the present invention are merely examples for clearly explaining the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, based on the above description, other different forms of changes or modifications can be made. It is impossible to list all the implementation manners here. Any obvious changes or modifications derived from the technical solutions of the present invention still fall within the protection scope of the present invention.
[0103] The content not detailed in this specification belongs to the prior art well-known to those of ordinary skill in the art.
Claims
1. An information cascade prediction method based on multi-task learning, characterized in that, It includes the following steps: S1. Preprocess the original information cascade, arrange the users participating in the propagation according to the forwarding time to obtain a forwarding sequence, and slice the cascade graph according to the time interval to obtain a snapshot sequence composed of snapshots at each slicing moment; S2. Initialize the feature representation of each user in the forwarding sequence obtained in step S1 using one-hot encoding, and input it into a long short-term memory network to obtain the corresponding microscopic information representation H user ; S3. Use the method based on the temporally improved bidirectional graph convolutional neural network to model the snapshot sequence obtained in step S1, and obtain the macroscopic information representation H containing cascade dynamic evolution information cas ; S4. Use the improved shared gating mechanism to fuse the above microscopic information representation H user and the macroscopic information representation H cas to obtain the shared representation H that fuses the information at both levels share ; The specific steps are as follows: Given the microscopic information representation H user and the macroscopic information representation H cas , the designed shared gating mechanism is used to aggregate the information at both levels to obtain the shared representation H share that contains all the information; the calculation process of the proposed shared gating mechanism is as follows: First, the forget gate f t is used to select the irrelevant information in the previous state to be forgotten and update the cell state c t , and the calculation method is shown in Equation (3): where σ(·) represents the sigmoid activation function, W f , U f and b f represent the learnable parameters during the training process of the gated neural network, and represent the microscopic information representation and macroscopic information representation at time t-1 respectively; Use the obtained forget gate to filter the information to be forgotten in the cell state at time t-1, and update to obtain a new cell state c t , and the calculation method is shown in Equation (4). W c and U c represent the learnable parameters in the training process of the gated neural network; Reset gate r t For determining the input of the new hidden layer state h t The information in it is calculated as shown in Equation (5), where W r , U r and b r represent the learnable parameters during the training process of the gated neural network; Use the obtained reset gate to perform information selection on the new cell state c t and update to obtain the new hidden layer state h t , and the calculation method is as shown in Equation (6). W h and U h represent the learnable parameters during the training process of the gated neural network; Fuse H using the above shared gating mechanism user and H cas and use the hidden state of the last step of the gated recurrent neural network as the shared representation layer H of multi-task learning share ; In the above formulas (3)-(6), σ(·) represents the sigmoid activation function, and ⊙ represents element-wise multiplication; S5. Concatenate the shared representation H obtained in step S4 share with the information representations corresponding to the respective tasks, input the micro-concatenated representation in the above information representation concatenation into the softmax function to predict the next user participating in the propagation, and input the macro-concatenated representation in the above information representation concatenation into a multi-layer perceptron to predict the popularity increment Δy.
2. The information cascade prediction method based on multi-task learning according to claim 1, characterized in that: The specific steps of step S1 are as follows: Given the original information cascade G, arrange the users participating in the propagation according to the forwarding time to obtain a user forwarding sequence S = {v1, v2,..., v N}, where v * represents each user, and N represents the number of users; divide the given observation window T into n time intervals, and count the propagation graph information at each slice moment to obtain a snapshot g i = {u i , e i , t i}, where u i , e i and t i represent the set of users, the set of edges, and the set of forwarding timestamps in the i-th snapshot respectively. Finally, obtain a snapshot sequence G = {g1, g2,..., g n}.
3. The information cascade prediction method based on multi-task learning according to claim 1, wherein: The specific steps of step S2 are as follows: Given the forwarding sequence S = {v1, v2,..., v N}, initialize the feature representation of each user using one-hot encoding; all users will be associated with a specific embedding matrix E ∈ R N×D , where D represents the embedding dimension of the vector; for user u, map it into the corresponding low-dimensional dense representation vector x = uE using the embedding matrix E; use the long short-term memory network LSTM to model the forwarding sequence after vector representation. Assume that the embedding vector of the user at time t is x i , and the hidden layer state at time t-1 is h i-1 . The calculation method of the hidden layer state at time t is shown in Equation (1): h i = LSTM(x i , h i-1 ) (1); Among them d seq represents the dimension of the hidden layer unit.
4. The information cascade prediction method based on multi-task learning according to claim 1, characterized in that: The specific steps of step S3 are as follows: Given a snapshot sequence \(G = \{g_1, g_2, \ldots, g\}\), for the \(i\)-th snapshot \(g\), n using the TGAT method to encode the timestamp set \(t\) in the snapshot, i a temporal matrix \(C\) is obtained; i Based on obtaining the adjacency matrix \(A\) of the cascaded graph corresponding to the \(i\)-th snapshot \(g\), i the degree matrix \(D\) in \(g\) and the previous temporal matrix \(C\) are respectively input into the Bi-directional Graph Convolutional Neural Network (BiGCN), i to obtain the corresponding structural information representation i and temporal information representation. After concatenating the two types of information representations, i the representation corresponding to each snapshot is obtained. The calculation method is shown in Equation (2): i Among them, I N is the identity matrix, and W C and W D represent trainable parameter matrices; The characterization of each snapshot can form a sequence of snapshot representations Input it into a gated neural network to extract the temporal dependencies between different snapshots, and obtain the characterization matrix H of the final cascaded graph cas .
5. The information cascade prediction method based on multi-task learning according to claim 1, characterized in that: The specific steps of step S5 are as follows: Given the shared representation layer H share , the specific representations H user and H cas of each task are respectively concatenated with it, and the concatenated specific task representations are input into different output layers for prediction; the goal of the micro-cascade prediction task is to predict the next user participating in the cascade propagation. The output layer uses the softmax function of the logical classifier for prediction. For each user, the calculation method of the probability that the model predicts that it will participate in the propagation at the i+1 moment is shown in Equation (7), where concat(·) represents the concatenation operation: Given the length l of the cascade sequence, the learning objective of the micro-cascade prediction model is to maximize the probability of all users participating in the propagation in the propagation sequence, and the objective function is shown in formula (8): The loss function of the micro-prediction task during model training can be regarded as minimizing the cross-entropy loss between the predicted probability of sequence S and the true probability P, and the calculation method is shown in Equation (9): The goal of the macro-cascade prediction task is to predict the popularity increment in a future fixed time interval by observing the cascade information within a fixed time window. The output layer uses a regressor multi-layer perceptron MLP for prediction. For the snapshot sequence G formed by the information cascade, the calculation method of the model to predict the popularity increment is shown in formula (10): ΔS = MLP(concat(h cas , h share )) (10); The loss function during model training for the macro-prediction task can be regarded as minimizing the error between the predicted value and the true value, and the calculation method is shown in formula (11): where Q represents the total number of cascades in the dataset, represents the true popularity increment; The loss function of the overall model is shown in Equation (12), where γ ∈ [0, 1] is the learning parameter that balances and :
Citation Information
Patent Citations
Social network information propagation scale prediction method and device
CN113536144A
Information diffusion prediction method fusing space-time attention and heterogeneous graph convolutional network
CN113850446A