A propagation source localization method based on encoder and decoder framework
Through the method based on the encoder and decoder framework, users' influence and dynamic characteristics are evaluated, combined with the timing attention mechanism, the problems of low accuracy and poor generalization of propagation source positioning in the prior art are solved, and efficient and accurate positioning and model migration in complex scenarios are achieved.
Patent Information
- Application Number
- CN202310557470.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2043-05-17
AI Technical Summary
The existing technology lacks a sequence-to-sequence framework to efficiently model time-varying infection data, and does not fully consider the impact of user heterogeneous influence on propagation. The trained model has poor generalization and is single applicable scenarios, resulting in low accuracy in propagation source positioning in complex scenarios.
Using an encoder and decoder framework method, sequence-to-sequence learning and inductive learning are achieved by evaluating the influence transfer matrix between users, dynamic infection characteristics and topological characteristics of constructing nodes, and combining bidirectional GRU and multi-head timing attention mechanisms.
It improves the accuracy of communication source positioning and model applicability in complex scenarios, can migrate and apply in different social networks and communication models, and enhances the accurate positioning ability of user-oriented complex behavior interaction scenarios in social networks.
Smart Images

Figure CN117034134B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cyberspace security information tracing technology, and in particular to a propagation source positioning method based on an encoder and decoder framework. Background Art
[0002] The rapid development of internet technology has fueled the growing popularity of social platforms and brought unprecedented, profound changes to human production and lifestyles. Social media software has gradually become the primary channel for netizens to access and disseminate information. Based on the complex social network relationships established within social media, information dissemination is characterized by rapid speed, strong interactivity, and wide reach. While this tremendous convenience in information exchange has promoted social and economic development, it has also created opportunities for the spread of harmful online information. For example, the spread of online viruses severely impacts user experience and damages users' property and information. This harmful online information significantly impacts national security, social stability, and the public interest. Therefore, timely and effective locating the source of transmission is crucial for mitigating the impact of harmful information and minimizing losses.
[0003] Scholars at home and abroad widely use attribution tracing methods based on propagation subgraphs and observation points. However, observation points are sometimes not pre-deployed in real-world areas where outbreaks are imminent, and observer deployment requires significant overhead, limiting their application scenarios. In contrast, propagation subgraph-based methods capture easily accessible or permissible snapshots of network infections at accessible timestamps, offering a wide range of applications and low cost. Consequently, these methods have been extensively studied. For example, IVGD and SL_VAE consider the dynamic characteristics of propagation before performing source inference. However, these methods are only applicable when the underlying propagation model and the complete propagation chain at all timestamps are known. Although Dong et al. designed a GCN-based source identification model with no heuristic information (GCNSI) to overcome the difficulty of acquiring prior knowledge, related methods only consider simple static features of individuals. Consequently, some research focuses on more features in the task of localizing the source of infection. The SIGN method extracts static features of individuals and combines them with a variant of a GNN to predict the source of infection. MCGNN considers the topological features of edges and designs a multi-channel graph neural network framework for source localization tasks. However, it does not consider, or only briefly considers, the impact of individual characteristics on propagation. Therefore, in snapshots with less heuristic information, only simple explicit features are used to infer the source of propagation, resulting in relatively low positioning accuracy in complex user interaction environments. In fact, some discrete and discontinuous snapshots can be obtained at a low cost. Analyzing snapshots with time-varying infection characteristics makes it possible to assess the potential heterogeneous influence between any two users. Therefore, one of the key challenges in current research based on propagation subgraph methods is to comprehensively evaluate user interaction behavior by considering snapshot information under discrete available time series features, ultimately improving the accuracy of propagation source localization in complex scenarios.
[0004] In summary, the problems and defects of the existing technical methods are as follows:
[0005] 1. Lack of a sequence-to-sequence framework for efficient modeling of time-varying infection data. Current technical methods lack a universal sequence-to-sequence localization framework for processing snapshot data with time series characteristics. This inability to effectively measure the imbalance of information at different timestamps results in low localization accuracy.
[0006] Second, the attribution process fails to consider the impact of heterogeneous user influence on communication. A set of time-varying communication snapshots contains rich interactive information, but existing technical methods do not fully consider the impact of heterogeneous user behavior on communication, causing positioning technology to degenerate into a centrality-based solution. As a result, the attribution accuracy is low in complex communication scenarios.
[0007] Third, trained models cannot be effectively transferred to other scenarios, resulting in poor generalization and a limited use case. Existing traceability technologies based on static topology (for example, the addition of new users or changes in the relationships of existing users within the same social network will render the model unusable) severely limit the model's applicable scenarios.
[0008] Based on this, the present invention designs a propagation source positioning method based on an encoder and decoder framework to solve the above problems. Summary of the Invention
[0009] In view of the above-mentioned shortcomings of the prior art, the present invention provides a propagation source positioning method based on an encoder and decoder framework.
[0010] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0011] A method for locating a propagation source based on an encoder and decoder framework includes the following steps:
[0012] S1. Input a set of time-varying propagation subgraphs; input a set of snapshots obtained in the target area
[0013] Among them, V is the user set of the target area, E is the edge set, and an edge indicates that two users know each other in the social network. Timestamp s j The observed status information of all users under
[0014] S2, evaluate the influence transfer matrix between any two users, according to The snapshot information in the evaluation of the user's influence transfer matrix W, where W ij Indicates user v i For user v j influence;
[0015] S3. Evaluate the coarse-grained probability characteristics of each node at timestamp s
[0016] S4. Construct the dynamic infection characteristics of each node at timestamp s and
[0017] S5. Construct the topological characteristics H of each node in the network G = (V, E) G ;
[0018] S6. Aggregate at timestamp s j The characteristics of all nodes
[0019] S7. Use bidirectional GRU to predict the propagation source of time series data at all timestamps;
[0020] S8. Use the multi-head temporal attention mechanism to correct the propagation source prediction probability of any node v at different timestamps;
[0021] S9, the group of nodes with the highest output probability is the predicted propagation source;
[0022]
[0023] Furthermore, in step S1, Represents node v i At timestamp s j The status information below is that harmful information has been received or infected. Represents node v i At timestamp s j The status information below means that no harmful information has been received or infected.
[0024] Furthermore, in step S2, the evaluation of the influence transfer matrix between any two users specifically includes: using the propagation model f θ To characterize the influence transfer matrix W:
[0025]
[0026] Among them, W ji Indicates neighbor user v j For user v i influence, Represents node v i The predicted state at time s+1, Represents node v i The actual snapshot state at timestamp s, Represents node v i At the real snapshot state of timestamp s, N(v i ) represents node v i neighbors.
[0027] Furthermore, in step S3, Represents node v i Coarse-grained probabilistic feature representation of ;
[0028]
[0029] in, It is a GCN module with a residual structure. is the GCN module, Y srepresents the true snapshot state of all nodes in the network at timestamp s, W1 is the learnable weight matrix in the GCN module, A is the adjacency matrix representation of the network G = (V, E), X is the feature matrix of the nodes in the network, X i is the characteristic vector of the i-th node in the network, I n is the n-order identity matrix, W is the influence transfer matrix, σ(x)=1 / (1+e -x ) is the activation function, is the coarse-grained source probability output feature of all nodes at timestamp s.
[0030] Furthermore, in step S4, and Represents node v i Dynamic infection characteristics and dynamic non-infection characteristics;
[0031]
[0032] in, Representative node v i The infected neighbor information characteristics, Representative node v i The non-infected neighbor information feature, N(v i ) represents node v i The neighbor set of v j Representative node v i The j-th neighbor of , W represents the influence transfer matrix.
[0033] Furthermore, in step S5, H G (v i ) has a dimension of
[0034]
[0035] Among them, W G is a learnable parameter matrix, A is the adjacency matrix representation of the network G = (V, E), I n is the n-th order identity matrix, is the degree-diagonally normalized matrix of A, H G is the output topological representation of all nodes.
[0036] Furthermore, in step S7, the transmission source task is a binary classification task, i.e., it is divided into transmission source probability and non-transmission source probability;
[0037]
[0038] Among them, GRU represents the bidirectional GRU module, and are the propagation source prediction vectors of all nodes at time j-1 and j, respectively, H j Timestamp s j The characteristics of all nodes.
[0039] Furthermore, in step S7, the following formula is used:
[0040]
[0041] Among them, W ir 、W hr 、W iz 、W hz 、W in 、W hn is the training parameter, b ir 、b hr 、b iz 、b hz 、b in 、b hn is the bias parameter, σ(x)=1 / (1+e -x ) is the activation function, tanh function is the nonlinear feature transformation, z and r are the update gate and forget gate respectively, n j Timestamp s j The state vector under is the propagation source prediction vector of all nodes at time j-1, H j Timestamp s j The characteristics of all nodes.
[0042] Furthermore, in step S8, the importance of node v at time j to time i is defined as follows:
[0043]
[0044] in, W A is a learnable parameter, is the propagation source prediction vector of node v at time j.
[0045] Beneficial effects
[0046] First, the method of the present invention is based on the framework of sequence-to-sequence learning. It designs an encoder that integrates user interaction behavior and a decoder based on temporal attention. This solves the problem of poor positioning accuracy caused by ignoring the impact of heterogeneous user behavioral interactions on propagation, and achieves accurate positioning for complex user behavioral interaction scenarios in social networks. In addition, the positioning model used in the present invention adopts the concept of inductive learning, so that the trained positioning model can be applied to different social networks or propagation models. This enhances the applicable scenarios of the model and makes the present invention highly transferable and universal in the field of traceability.
[0047] Second, this method requires no prior knowledge and manually designs coarse-grained propagation source probability features, dynamic infection features, and static topological features for each user. This enhances the expressiveness of the user representation vector, thereby increasing the model's positioning accuracy in complex scenarios. Furthermore, a one-step temporal attention mechanism is designed to distinguish the importance of information at different timestamps. Experiments show that this method's positioning accuracy outperforms other state-of-the-art methods.
[0048] 3. The expected benefits and commercial value of the technical solution of the present invention after transformation are as follows: the use of the present invention to perform the task of locating the source of transmission under the security of cyberspace can play a key role in applications such as tracing the source of social media rumors and tracking the source of network viruses; specifically, digging out and tracking the initiators of rumors or viruses plays a vital role in cutting off the transmission channels, stopping losses in time, stifling malicious information, and thus purifying the network environment and maintaining social stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0050] Figure 1 Flowchart of the method for locating a propagation source based on an encoder and decoder framework of the present invention;
[0051] Figure 2 The principle of the transmission source positioning method based on the encoder and decoder framework of the present invention Figure 1 ;
[0052] Figure 3 The principle of the transmission source positioning method based on the encoder and decoder framework of the present invention Figure 2 ;
[0053] Figure 4 This is a comparison chart of the network migration performance evaluation of the propagation source localization method based on the encoder and decoder framework of the present invention and other methods in six real networks. DETAILED DESCRIPTION
[0054] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0055] The present invention will be further described below with reference to the embodiments.
[0056] Example 1
[0057] Figure 2 The principle of the transmission source positioning method based on the encoder and decoder framework of the present invention Figure 1 The present invention aims to solve the problem of detecting multiple transmission sources based on discrete available timestamps. Since obtaining a complete transmission chain at all times requires the equipment to have high-precision statistical and capture capabilities, it is not easy to achieve. In contrast, randomly and non-intentionally collecting infection snapshots in social networks at a few accessible timestamps is less costly and more feasible. Specifically, a transmission starts at t0 and goes through six stages, t1, t2, t3, t4, and t5, to achieve information coverage of the entire network. The embodiment of the present invention randomly captures the network snapshot infection information at timestamps t2 and t4. Therefore, based only on the information at these two snapshots, the present invention locates the transmission source. At the same time, the snapshot information at t2 and t4 is also the input of the method of the present invention.
[0058] Figure 1 Flowchart of the method for locating a propagation source based on an encoder and decoder framework of the present invention; Figure 3 The principle of the transmission source positioning method based on the encoder and decoder framework of the present invention Figure 2 ;
[0059] A method for locating a propagation source based on an encoder and decoder framework includes the following steps:
[0060] S1, input a set of time-varying propagation subgraphs;
[0061] Enter a set of snapshots obtained in the target area
[0062] Among them, V is the user set of the target area, E is the edge set, and an edge indicates that two users know each other in the social network. Timestamp s j The observed status information of all users under
[0063] Preferably, in step S1, Represents node v i At timestamp s j The status information below is that harmful information has been received or infected. Represents node v i At timestamp s j The status information below is that no harmful information has been received or infected;
[0064] S2, evaluate the influence transfer matrix between any two users ( Figure 3 (a)), we get a matrix |V|*|V|, according to The snapshot information in the evaluation of the user's influence transfer matrix W, where W ij Indicates user v i For user v j influence;
[0065] Preferably, in step S2, the evaluation of the influence transfer matrix between any two users specifically includes: using the propagation model f θ To characterize the influence transfer matrix W:
[0066]
[0067] Among them, W ji Indicates neighbor user v j For user v i influence, Represents node v i The predicted state at time s+1, Represents node v i The actual snapshot state at timestamp s, Represents node v i At the real snapshot state of timestamp s, N(v i ) represents node v i neighbors; use the real snapshot information to guide the prediction of the state of each node in the next time step, and finally realize the learning of the influence transfer matrix W;
[0068] S3. Evaluate the coarse-grained probability characteristics of each node at timestamp s ( Figure 3 (b)), we get a vector of length |V|, where Represents node v i Coarse-grained probabilistic feature representation of ;
[0069]
[0070] in, It is a GCN module with a residual structure. is the GCN module, Y s represents the true snapshot state of all nodes in the network at timestamp s, W1 is the learnable weight matrix in the GCN module, A is the adjacency matrix representation of the network G = (V, E), X is the feature matrix of the nodes in the network, X i is the characteristic vector of the i-th node in the network, I n is the n-order identity matrix, W is the influence transfer matrix, σ(x)=1 / (1+e -x ) is the activation function, is the coarse-grained source probability output feature of all nodes at timestamp s;
[0071] S4. Construct the dynamic infection characteristics of each node at timestamp s and ( Figure 3 (c)), we get two vectors of length |V|, where and Represents node v i Dynamic infection characteristics and dynamic non-infection characteristics;
[0072]
[0073] in, Representative node v i The infected neighbor information characteristics, Representative node v i The non-infected neighbor information feature, N(v i ) represents node v i The neighbor set of v j Representative node v i The j-th neighbor of , W represents the influence transfer matrix;
[0074] S5. Construct the topological characteristics H of each node in the network G = (V, E) G ( Figure 3 (d));
[0075] Among them, H G (v i ) has a dimension of (Round up), taking the dataset G1 used in the embodiment as an example, there are 198 nodes, then |V| is 198, Round up to 15, then H G The dimension is (198, 15), H G (v i ) dimension is 15;
[0076]
[0077] Among them, W G is a learnable parameter matrix, A is the adjacency matrix representation of the network G = (V, E), I n is the n-th order identity matrix, is the degree-diagonally normalized matrix of A, H G is the output topological representation of all nodes;
[0078] S6. Aggregate at timestamp s j The characteristics of all nodes
[0079] Taking the dataset G1 used in the embodiment as an example, there are 198 nodes, so |V| is 198, and after aggregation H G The dimension is (198, 18);
[0080] S7, use bidirectional GRU to predict the source of transmission of time series data at all timestamps ( Figure 3 (e));
[0081] Among them, the propagation source task is a binary classification task, that is, it is divided into the probability of propagation source and the probability of non-propagation source. Therefore, the output dimension of the unidirectional GRU is (198, 2), and the output dimension of the bidirectional GRU is (198, 4);
[0082]
[0083] Among them, GRU represents the bidirectional GRU module, and are the propagation source prediction vectors of all nodes at time j-1 and j, respectively, H j Timestamp s j The characteristics of all nodes below;
[0084] In step S7, the following formula is used:
[0085]
[0086] Among them, W ir 、W hr 、W iz 、W hz 、W in 、W hn is the training parameter, b ir 、b hr 、b iz 、b hz 、b in 、b hn is the bias parameter, σ(x)=1 / (1+e -x ) is the activation function, tanh function is the nonlinear feature transformation, z and r are the update gate and forget gate respectively, n j Timestamp s j The state vector under is the propagation source prediction vector of all nodes at time j-1, H j Timestamp s j The characteristics of all nodes below;
[0087] S8. Use the multi-head temporal attention mechanism to correct the propagation source prediction probability of any node v at different timestamps;
[0088] Taking the data set G1 used in the embodiment as an example, there are 198 nodes, and the propagation source task is a binary classification task, that is, it is divided into propagation source probability and non-propagation source probability, so the final output is (198, 2);
[0089]
[0090] Where K is the number of attention heads, σ(x)=1 / (1+e -x ) is the activation function, W A is a learnable parameter, is the propagation source prediction vector of all nodes at time j, is the corrected propagation source prediction vector of all nodes at time j. The importance of node v at time j to time i is defined as follows:
[0091]
[0092] in, W A is a learnable parameter, is the propagation source prediction vector of node v at time j;
[0093] S9, the group of nodes with the highest output probability is the predicted propagation source;
[0094]
[0095] Among them, R * represents a group of nodes that are judged to be the most likely transmission sources, and Z is the number of predicted transmission sources.
[0096] Experimental example
[0097] The embodiments of the present invention have achieved some positive results during the development or use process, and indeed have great advantages over the existing technology. The following content describes them in conjunction with data, charts, etc. from the experimental process.
[0098] Table 1 shows the size of the test dataset; Table 1 shows the network structure information used in the test. G1-G6 are all real social network datasets;
[0099] Table 1. The size of the test dataset
[0100]
[0101] Table 2 shows the accuracy comparison results of the proposed method and other methods on six real networks;
[0102] Table 2 Comparison of the accuracy evaluation results of the present invention with other methods on six real networks
[0103]
[0104] 10% of the nodes in the corresponding six data sets are selected as propagation sources, and then an authoritative independent cascade model is used to model the propagation process. Then, 5 time steps are randomly selected to represent the easily accessible network infection snapshots obtained by the information capture device in reality at the accessible moment. The embodiment of the present invention generates 1000 groups in each data set, selects 90% as the test set, and 10% as the validation set. "F-score" (F1) and average error distance (AED) are used as evaluation indicators in each network. The complete parameters in the table are explained as follows: the entire row G1 to G6 corresponding to Network is the data set used in this embodiment. The effectiveness of the algorithm is proved by testing on six data sets of different sizes. The Algorithm column represents the authoritative comparative algorithm for propagation source positioning in recent years. Two evaluation indicators F1 and AED, where F1 is the prediction accuracy evaluation indicator. The higher the indicator, the stronger the algorithm's ability to predict the true source. AED represents the average value of the minimum offset distance between the true source and the predicted source, that is, the error distance. The smaller the indicator, the stronger the algorithm's ability to predict the true source.
[0105] Figure 4 This is a comparison chart of the transfer ability of the time series-based graph attention source identification method provided by the embodiment of the present invention and other comparison methods on six real networks. Figure 4 It can be seen that the proposed method for localizing the source of transmission based on the encoder and decoder framework (TGASI) outperforms other methods on all networks;
[0106] The embodiment of the present invention applies the propagation source positioning method based on the encoder and decoder framework trained in the G3 network to the propagation data of other networks G1-G6 (excluding G3) to predict the propagation source (such as Figure 4 (a) shows), it can be seen that compared with the prediction results of the best comparison algorithm on their respective original network datasets (such as Figure 4 (b) shows that the migration prediction ability of the source localization method based on the encoder and decoder framework is higher than the original prediction ability of the comparison method, which shows that the method of the present invention has a strong inductive learning ability;
[0107] The method of the present invention evaluates the influence transfer matrix between users through a deep module based on random walk stationary distribution, and then distinguishes the impact of users' complex behavioral characteristics on propagation; further, the present invention designs dynamic infection features and static topological features to enrich the representation of low-dimensional embedding; the dynamic features better fit the task of locating the propagation source, while the topological features enhance the learning ability of the decoder based on sequence tasks on graph tasks; finally, the present invention introduces a temporal attention mechanism to further comprehensively consider the weight of the source prediction probability at different timestamps, and solves the problem of imbalanced infection information at different time steps by dynamically distinguishing the importance of different timestamps; 6 real social networks are used to show that the algorithm proposed by the present invention solves the problem of low positioning accuracy when any prior information other than observation information is unavailable in a complex propagation environment; at the same time, a sequence-to-sequence propagation source positioning framework is designed in an inductive learning manner, so that the model does not require any prior knowledge and can be applied to new diffusion models and new social networks, making the method of the present invention have practical guiding significance;
[0108] The present invention has the following technical effects:
[0109] First, the method of the present invention is based on the framework of sequence-to-sequence learning. It designs an encoder that integrates user interaction behavior and a decoder based on temporal attention. This solves the problem of poor positioning accuracy caused by ignoring the impact of heterogeneous user behavioral interactions on propagation, and achieves accurate positioning for complex user behavioral interaction scenarios in social networks. In addition, the positioning model used in the present invention adopts the concept of inductive learning, so that the trained positioning model can be applied to different social networks or propagation models. This enhances the applicable scenarios of the model and makes the present invention highly transferable and universal in the field of traceability.
[0110] Second, this method requires no prior knowledge and manually designs coarse-grained propagation source probability features, dynamic infection features, and static topological features for each user. This enhances the expressiveness of the user representation vector, thereby increasing the model's positioning accuracy in complex scenarios. Furthermore, a one-step temporal attention mechanism is designed to distinguish the importance of information at different timestamps. Experiments show that this method's positioning accuracy outperforms other state-of-the-art methods.
[0111] 3. The expected benefits and commercial value of the technical solution of the present invention after transformation are as follows: the use of the present invention to perform the task of locating the source of transmission under the security of cyberspace can play a key role in applications such as tracing the source of social media rumors and tracking the source of network viruses; specifically, digging out and tracking the initiators of rumors or viruses plays a vital role in cutting off the transmission channels, stopping losses in time, stifling malicious information, and thus purifying the network environment and maintaining social stability.
[0112] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A propagation source localization method based on an encoder and decoder framework, characterized in that: The following steps are involved: S1. Input a set of time-varying propagation subgraphs; input a set of snapshots obtained in the target area Among them, V is the user set of the target area, E is the edge set, and an edge indicates that two users know each other in the social network. Timestamp s j The observed status information of all users under Represents node v i At timestamp s j The status information below is that harmful information has been received or infected. Represents node v i At timestamp s j The status information below is that no harmful information has been received or infected; S2, evaluate the influence transfer matrix between any two users, according to The snapshot information in the evaluation of the user's influence transfer matrix W, where W ij Indicates user v i For user v j influence; S3. Evaluate the coarse-grained probability characteristics of each node at timestamp s Represents node v i Coarse-grained probabilistic feature representation of ; in, It is a GCN module with a residual structure. is the GCN module, Y s represents the true snapshot state of all nodes in the network at timestamp s, W1 is the learnable weight matrix in the GCN module, A is the adjacency matrix representation of the network G = (V, E), X is the feature matrix of the nodes in the network, X i is the characteristic vector of the i-th node in the network, I n is the n-order identity matrix, W is the influence transfer matrix, σ(x)=1 / (1+e -x ) is the activation function, is the coarse-grained source probability output feature of all nodes at timestamp s; S4. Construct the dynamic infection characteristics of each node at timestamp s and and Represents node v i Dynamic infection characteristics and dynamic non-infection characteristics; in, Representative node v i The infected neighbor information characteristics, Representative node v i The non-infected neighbor information feature, N(v i ) represents node v i The neighbor set of v j Representative node v i The j-th neighbor of , W represents the influence transfer matrix; S5. Construct the topological characteristics H of each node in the network G = (V, E) G ; S6. Aggregate at timestamp s j The characteristics of all nodes S7. Use bidirectional GRU to predict the propagation source of time series data at all timestamps; S8. Use the multi-head temporal attention mechanism to correct the propagation source prediction probability of any node v at different timestamps; S9, the group of nodes with the highest output probability is the predicted propagation source; Among them, R * represents a group of nodes that are judged to have the highest probability of being the source of transmission, and Z is the number of predicted transmission sources. is the corrected propagation source prediction vector of all nodes at time j.
2. The method for locating a propagation source based on an encoder and decoder framework according to claim 1, characterized in that: In step S2, the influence transfer matrix between any two users is evaluated, specifically including: using the propagation model f θ To characterize the influence transfer matrix W: Among them, W ji Indicates neighbor user v j For user v i influence, Represents node v i The predicted state at time s+1, Represents node v i The actual snapshot state at timestamp s, Represents node v i At the real snapshot state of timestamp s, N(v i ) represents node v i neighbors.
3. The method for locating a propagation source based on an encoder and decoder framework according to claim 1, wherein: In step S5, H G (v i ) has a dimension of Among them, W G is a learnable parameter matrix, A is the adjacency matrix representation of the network G = (V, E), I n is the n-th order identity matrix, is the degree-diagonally normalized matrix of A, H G is the output topological representation of all nodes.
4. The method for locating a propagation source based on an encoder and decoder framework according to claim 1, wherein: In step S7, the transmission source task is a binary classification task, that is, it is divided into the probability of transmission source and the probability of non-transmission source; Among them, GRU represents the bidirectional GRU module, and are the propagation source prediction vectors of all nodes at time j-1 and j, respectively, H j Timestamp s j The characteristics of all nodes.
5. The method for locating a propagation source based on an encoder and decoder framework according to claim 4, wherein: In step S7, the following formula is used: Among them, W ir 、W hr 、W iz 、W hz 、W in 、W hn is the training parameter, b ir 、b hr 、b iz 、b hz 、b in 、b hn is the bias parameter, σ(x)=1 / (1+e -x ) is the activation function, tanh function is the nonlinear feature transformation, z and r are the update gate and forget gate respectively, n j timestamp s j The state vector under is the propagation source prediction vector of all nodes at time j-1, H j Timestamp s j The characteristics of all nodes.
6. The method for locating a propagation source based on an encoder and decoder framework according to claim 1, wherein: In step S8, the importance of node v at time j to time i is defined as follows: in, W A is a learnable parameter, is the propagation source prediction vector of node v at time j.
Citation Information
Patent Citations
Malicious information tracing method based on neighborhood similarity and multi-type interaction
CN115375328A
Network topology reconstruction method and apparatus, and terminal device
WO2020093291A1