A multi-source localization method based on deep learning

Through a multi-source positioning method based on deep learning, using self-coded networks and epidemic transmission SI model, the problems of large amount of computing and low efficiency in the prior art are solved, and efficient multi-source positioning in social networks are achieved.

CN115018662BActive Publication Date: 2025-05-23YANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210658016.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2025-05-23
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

The existing multi-source positioning method has a large amount of computation in social networks, resulting in long detection time and low efficiency.

Method used

The multi-source positioning method based on deep learning is adopted to compile, integrate and decode node features through a self-encoding network, and spread the network with the epidemic transmission SI model, generate an infection sub-graph and extract node features, and finally determine the source node through probability calculation.

Benefits of technology

It improves the operation efficiency of the algorithm, reduces the computational complexity and investment cost, and can achieve better accuracy and small errors when obtaining some infection information, which significantly improves the accuracy of early source positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115018662B_ABST
    Figure CN115018662B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of social network source positioning, and specifically to a multi-source positioning method based on deep learning, which integrates node features in combination with an auto-encoder network (AE), and uses the advantages of deep learning for large amounts of data to improve the overall operating efficiency of the algorithm. The present invention combines the relative relationship between time and distance in a graph to derive a method for extracting node features, comprehensively considers the possibility of node propagation paths and time conditions, and has a more detailed description of node features, retaining most of the propagation information and properties in the infected subgraph. This method enables the algorithm to achieve relatively good accuracy and small errors when only a portion of the infection information is obtained and a small number of observation nodes are extracted, greatly reducing the input cost and computational complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of social network source positioning, and in particular to a multi-source positioning method based on deep learning. Background Art

[0002] Research on accurately locating the source of false information dissemination in a short period of time is of great practical significance and application value.

[0003] Based on existing research, single-source problems are more widely studied, while multi-source problems are often accompanied by relatively complex propagation situations, chaotic temporal information, and complex infection group graph information. Some algorithms can only obtain good accuracy and low errors in tree networks or tree-like networks. Moreover, fewer algorithms are based on neural networks and deep learning, which require a large amount of computation, resulting in longer detection time and lower efficiency on medium and large social networks. Summary of the invention

[0004] In view of this, the purpose of the present invention is to propose a multi-source positioning method based on deep learning to solve the problem that the existing methods have large computational complexity, resulting in long detection time and low efficiency.

[0005] Based on the above purpose, the present invention provides a multi-source positioning method based on deep learning, comprising the following steps:

[0006] Randomly select k propagation source nodes s* in the original undirected network;

[0007] Based on the epidemic propagation SI model, the original network is diffused until the round ends or no new infected nodes are generated. Then the nodes and edges generated after the diffusion of the source nodes are extracted to generate the infected subgraph G infected ;

[0008] According to the infected subgraph, a certain proportion of infected nodes are randomly selected as the observer node set O = {o 1 , o 2 , ..., o m}, record the infection time of the selected node and obtain the infection time set T;

[0009] Construct node feature extraction: Based on random walk, obtain k paths of a certain length of infected nodes, and take the union to obtain the path of any observed node o. i The associated node set C i , extract C i The associated nodes in are relative to the observation node o i The distance parameter α, and according to the set C i Calculate each node v with the set O i The vector set X of the characteristics vi ={x v 1 , x v 2 …x v i …};

[0010] Then, by sorting, we can get the v node relative to each observer o i Vector set X v ;

[0011] Constructing a self-encoding network framework for v i The original feature values ​​are continuously compiled, integrated, and decoded, and the loss of each round of iteration is calculated to generate the integrated feature z i ;

[0012] Combined with the generated integrated features z i Perform probability calculation to obtain the likelihood evaluation score of the node as the source, and select the first k nodes as the estimated source set.

[0013] Preferably, constructing node feature extraction specifically includes: first considering the node propagation path, and i The node simulates the infection path and obtains k paths l of a certain length through random walk, which can be expressed as Generate path set C i , expressed as Extract the distance parameter α of all nodes relative to the observation node, denoted as α represents the node v i to v j The spatiotemporal relationship between nodes is recorded, and the distance between nodes is obtained, where α represents the node v i to v j time-space relationship;

[0014] According to the set C i and the set O computes the set of nodes V = {v 1 , v 2 , …, v k The characteristic vector group x of each node v in} v , expressed as

[0015] Preferably, The calculation formula conforms to:

[0016]

[0017] where t i is the infection time corresponding to the i-th node in the infection time set T.

[0018] Preferably, construct an autoencoding network framework for v i The original feature values ​​are continuously compiled, integrated, and decoded, and the loss of each round of iteration is calculated to generate the final integrated feature z i include:

[0019] After the node feature extraction, the overall information of all nodes enters the multi-layer autoencoder network for training, and the encoding, integration, and decoding processes are repeated until the number of iterations is met or the loss reaches the threshold;

[0020] The overall autoencoder network consists of two parts: the encoder and the decoder. The two layers are connected by the node integration features required by the positioning algorithm, that is, the integration layer. At the same time, the loss function of the autoencoder network is defined from four aspects: encoding loss, α difference loss, time difference loss and regularization term, to generate the final integrated feature z i .

[0021] Preferably, the method further comprises:

[0022] Constructing a self-encoding network framework for v i The original feature value is compiled, and the encoding layer is To iterate, the decoding layer iterates with the following formula group:

[0023]

[0024]

[0025]

[0026]

[0027] Calculate the loss function obtained in each round of iteration to control the number of iterations. The loss function consists of four parts: Loss of the x encoding x ; Code z n With z v The difference between uv Matching loss Loss α ; Node v and its C i Medium i Time matching loss o ; Regularization term L reg ; The final loss is the sum of the four losses.

[0028] Preferably, the calculation formula of the comprehensive loss Loss is:

[0029] Loss=Loss x +αLoss α +βLoss o+γLoss reg

[0030]

[0031]

[0032]

[0033]

[0034] Preferably, the integrated feature z is generated i The formula is:

[0035]

[0036]

[0037] Preferably, combined with the generated integrated features z i The formula for probability calculation is:

[0038]

[0039] Among them, w and b are weights and offsets, which are randomly generated at the beginning of the neural network and gradually determined as the training progresses; y v is the transfer vector of the v node vector in the hidden layer of the encoding layer, is the transfer vector in the Lth hidden layer of the encoding layer, α, β and γ are hyperparameters preset by the experiment, z is the output result of the network integration layer, and m is the number of hidden layers of the neural network.

[0040] Preferably, the diffusion of the original network based on the epidemic propagation SI model specifically includes:

[0041] Change the status of all nodes in the set to I state, and record the infection time as t 0 , enter the diffusion process;

[0042] Set the threshold probability p that a node can be successfully infected infected In each round, the node marked as I state propagates to all its neighbors with its own probability, that is, the edge weight. If the threshold is met, the propagation is successful, and the state of the infected node changes from S state to I state. The above process continues until the end of the round or no new nodes are infected;

[0043] After the infection process is completed, the original graph is pruned to extract the infected nodes and related edges to obtain the infected subgraph G infected .

[0044] Preferably, the proportion of infected nodes selected as the observer node set O is 5%.

[0045] Beneficial effects of the present invention: The present invention proposes a deep learning framework different from the traditional source localization algorithm, combines the auto-encoder network (Auto-Encoder, AE) to integrate node features, and uses the advantages of deep learning for large amounts of data to improve the overall operating efficiency of the algorithm;

[0046] The present invention combines the relative relationship between time and distance in the graph to derive a method for extracting node features, comprehensively considers the node's propagation path possibility and time conditions, and provides a more detailed description of node features, retaining most of the propagation information and properties in the infection subgraph. This method enables the algorithm to achieve relatively good accuracy and small errors when only a portion of the infection information is obtained and a small number of observation nodes are extracted, greatly reducing the input cost and computational complexity.

[0047] This method has been verified multiple times in two artificial networks and four medium-to-large real networks. The results show that compared with other traditional source localization algorithms, this method performs more outstandingly and is more helpful for early source localization problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0049] Figure 1 This is a schematic diagram of the overall algorithm flow of the traceability framework of this method;

[0050] Figure 2 Comparison of the accuracy of different algorithms (RC, JC, DMP, EPA, DL-SIA) on different networks (BA, WS, Email-Eu-core, ego Facebook, US Power Grid, Wiki-Vote) in the single-source case;

[0051] Figure 3 Comparison of average error distances of different algorithms (RC, JC, DMP, EPA, DL-SIA) on different networks (BA, WS, Email-Eu-core, ego Facebook, US Power Grid, Wiki-Vote) in the single-source case;

[0052] Figure 4Comparison of average error distances of different algorithms (Net Sleuth, K-Center, DL-SIA) on different networks (BA, WS, Email-Eu-core, ego Facebook, US Power Grid, Wiki-Vote) when the number of sources is 2;

[0053] Figure 5 Comparison of average error distances of different algorithms (Net Sleuth, K-Center, DL-SIA) on different networks (BA, WS, Email-Eu-core, ego Facebook, US Power Grid, Wiki-Vote) when the number of sources is 5;

[0054] Figure 6 The impact of three loss calculation methods with different proportions on the total loss reduction rate in different networks (Email-Eu-core, ego Facebook, US Power Grid, Wiki-Vote) when the number of sources is 2;

[0055] Figure 7 The influence of different proportions of the number of observation nodes and the average length of the random walk generated paths on the average error distance when the source is 2 on different networks (BA, WS, Email-Eu-core, ego Facebook, USPower Grid, Wiki-Vote);

[0056] Figure 8 Comparison of the average running time of different algorithms (RC, JC, DMP, EPA, DL-SIA) on different networks (BA, WS, Email-Eu-core, ego Facebook, US Power Grid, Wiki-Vote) in the single-source case. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments.

[0058] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0059] like Figure 1-Figure 8 As shown, the embodiment of this specification provides a multi-source positioning method based on deep learning, including the following steps:

[0060] S101 randomly selects k propagation source nodes s* that meet the experimental requirements in the original undirected network;

[0061] Based on the epidemic propagation SI model, the original network is diffused. In each infection round or moment, any infected node has a certain probability of infecting its neighboring nodes until the round ends or no new infected nodes are generated. Then, the nodes and edges generated after the source node is diffused are extracted to generate the infected subgraph G. infected ;

[0062] Specifically, the states of all nodes in the set are changed to state I, and the infection time is recorded as t 0 , enter the diffusion process;

[0063] Set the threshold probability p that a node can be successfully infected infected In each round, the nodes marked as I state will propagate to all their neighbors with their own probability, that is, the edge weight. If the threshold is met, the propagation is successful, and the infected node state changes from S state to I state. The above process continues until the end of the round or no new nodes are infected. After the infection process is completed, the original graph is pruned to extract the infected nodes and related edges to obtain the infected subgraph G infected .

[0064] According to the infected subgraph, a certain proportion of infected nodes are randomly selected as the observer node set O = {o 1 , o 2 , ..., o m}, record the infection time of the selected nodes, and obtain the infection time set T. For example, the selection ratio can be set to 5%.

[0065] Construct node feature extraction: Based on random walk, obtain k paths of a certain length of infected nodes, and take the union to obtain the path of any observed node o. i The associated node set C i , extract C i The associated nodes in are relative to the observation node o i The distance parameter α, and according to the set C i Calculate each node v with the set O i The vector set X of the characteristics v i ={x v 1 , x v 2 …x v i …};

[0066] Specifically, the feature extraction of building nodes includes: firstly, considering the node propagation path, and then i The node simulates the infection path and obtains k paths l of a certain length through random walk, which can be expressed as Generate path set C i , expressed as Extract the distance parameter α of all nodes relative to the observation node, denoted as α represents the node v i to v j The spatiotemporal relationship between nodes is recorded, and the distance between nodes is obtained, where α represents the node v i to v j time-space relationship;

[0067] According to the set C i and the set O computes the set of nodes V = {v 1 , v 2 , …, v k The characteristic vector group x of each node v in} v , expressed as

[0068] The calculation formula conforms to:

[0069]

[0070] where t i is the infection time corresponding to the i-th node in the infection time set T.

[0071] Then, by sorting, we can get the v node relative to each observer o i Vector set X v ;

[0072] Constructing a self-encoding network framework for v i The original feature values ​​are continuously compiled, integrated, and decoded, and the loss of each round of iteration is calculated to generate the integrated feature z i ;

[0073] Specifically, it includes:

[0074] After the node feature extraction, the overall information of all nodes enters the multi-layer autoencoder network for training, and the encoding, integration, and decoding processes are repeated until the number of iterations is met or the loss reaches the threshold;

[0075] The overall autoencoder network consists of two parts: the encoder and the decoder. The two layers are connected by the node integration features required by the positioning algorithm, that is, the integration layer. At the same time, the loss function of the autoencoder network is defined from four aspects: encoding loss, α difference loss, time difference loss and regularization term, to generate the final integrated feature z i .

[0076] Generate integrated feature z i The formula is:

[0077]

[0078]

[0079] The method further comprises:

[0080] Constructing a self-encoding network framework for v i The original feature value is compiled, and the encoding layer is To iterate, the decoding layer iterates with the following formula group:

[0081]

[0082]

[0083]

[0084]

[0085] Calculate the loss function obtained in each round of iteration to control the number of iterations. The loss function consists of four parts: Loss of the x encoding x ; Code z n With z v The difference between uv Matching loss Loss α ; Node v and its C i Medium i Time matching losso ; Regularization term L reg ; The final loss is the sum of the four losses.

[0086] The calculation formula of the above comprehensive loss Loss is:

[0087] Loss=Loss x +αLoss α +βLoss o +γLoss reg

[0088]

[0089]

[0090]

[0091]

[0092] Among them, w and b are weights and offsets, which are randomly generated at the beginning of the neural network and gradually determined as the training progresses; y v is the transfer vector of the v node vector in the hidden layer of the encoding layer, is the transfer vector in the Lth hidden layer of the encoding layer, α, β and γ are hyperparameters preset by the experiment, z is the output result of the network integration layer, and m is the number of hidden layers of the neural network.

[0093] Combined with the generated integrated features z i Perform probability calculation to obtain the likelihood evaluation score of the node as the source, and select the first k nodes as the estimated source set.

[0094] All the above steps will be repeated 100 times to reduce the impact of factors such as accidental errors on the experimental results. All the comparison algorithms involved in the experiment are based on the same SI diffusion model and are also repeated 100 times to improve the referenceability of the results.

[0095] Figure 2This is the result of the accuracy of single-source detection. The horizontal axis is the proportion of observation nodes selected, and the vertical axis is the accuracy. The results are obtained by relatively balanced values. Compared with traditional geocentric algorithms and path reasoning algorithms, DL-SIA (Deep-Learning-based Sources Identification Algorithm) based on deep learning has obvious advantages, especially in the ego Facebook network and Wiki-Vote network with relatively obvious social network characteristics, and the accuracy is improved by 25.0%-41.8% compared with other algorithms. Combining the curve trends of the six figures, it can be concluded that when the proportion of observer nodes reaches about 40%, the speed of DL-SIA algorithm accuracy improvement slows down and is relatively stable. This also shows that the DL-SIA algorithm can accurately locate the source of erroneous information under the premise of less network information, making the positioning work more economical and efficient.

[0096] Figure 3 The average error distance of each algorithm in two virtual networks (BA, SW) and four real networks (Email-bu Eu-core, ego Facebook, US Power Grid, Wiki-Vote) is shown. Figure 1 The structure is the same. It can be seen that the improvement of the DL-SIA algorithm is very significant compared with the centrality algorithms RC and JC. Taking the BA, SW and Email networks with a scale of about 1,000 nodes as an example, DL-SIA can reduce the predicted source distance from the real source by up to 2.18 hops compared with RC and JC, while maintaining a low proportion of observer nodes. In the larger Wiki-Vote network, the predicted source node can be reduced by about 4.82 hops compared with the real source node. In the early stage of detection, the average error distance of DL-SIA and the average error distance of the DMP algorithm are relatively close, and the advantage is not obvious. This is probably because the DMP algorithm originally needs to consider the transformation of the three states of S, I, and R, while in this experiment only the transformation probability of S and I needs to be considered. After the infection model is known, the centralized fixed parameters of the DMP algorithm are easier to derive, and when the proportion of observation nodes is low, the detection effect of the DMP algorithm is still relatively stable in the early stage when only part of the infection information is covered.

[0097] Figure 4The figure shows the comparison between the average error distance of the DL-SIA algorithm and the other two multi-source algorithms when the number of sources is 2. The horizontal axis is the error distance, and the vertical axis is the frequency of the error distance in the experiment. It can be seen that when the number of sources is 2, the difference between Net Sleuth, K-Center and DL-SIA algorithms is relatively not obvious, but combined with the peak of each algorithm in the figure, the DL-SIA algorithm is mostly concentrated around 0-2 hops, followed by K-Center, which is concentrated around 1-3 hops, while Net Sleuth performs poorly, concentrated around 2-5 hops. This is related to the strategy of the Net Sleuth algorithm to select multiple source nodes. In the above two algorithms, there is no process of eliminating the estimated source and then calculating the next source node. The Net Sleuth algorithm will choose to find the most likely node i, then eliminate the node, repeat the algorithm deduction, and obtain the most likely node after removing node i, which is likely to lose some key information in the figure.

[0098] Figure 5 When the number of source nodes is given as 5, it can be seen that although K-Center is not much different from this method when the number of source nodes is small, the advantage of the DL-SIA algorithm is relatively more obvious when the number of sources increases, and its average error distance can still be basically stabilized within 3 hops. For the Wiki-Vote network with larger nodes, the error distance of DL-SIA is still relatively concentrated around 1-3 hops. Compared with the K-Center and Net Sleuth algorithms, this method has an improvement of 22.9%-38.8%, and the average error distance has increased by nearly 50% at the highest.

[0099] Figure 6 The influence of the proportion of the three kinds of losses on the rate of decrease of the total loss is given. The following conclusions can be drawn from the legend: z On any data set, considering the feature value extracted from each node i and α i The difference in loss is not strongly correlated with the total loss. This conclusion can be seen from the contour chromatogram distribution in (a), (b), (c), and (d). From the ordinate, as the ordinate increases, the rate of decrease of the neural network loss does not increase significantly. Combined with the three-dimensional graph, we can roughly conclude that the value of α can be taken as much as possible in the range of 0.5-0.6 to obtain a faster rate of decrease. Compared with the insensitivity of the loss function to α, the value of β has a more obvious impact. β measures the node i in the detected path set C. iThe difference between the infected time and the inferred eigenvalue. The value of β is different for different networks, but most of them are concentrated in the range of 0.5-0.7. The values ​​on US Power Grid are relatively different. The experiment achieved a faster decline rate when β was around 0.2. The reason is that compared with other networks, the average path length of the Power Grid network is the lowest, and its average clustering coefficient is low. At the same time, although the Power Grid contains 4941 nodes, the number of edges is only 6594. It simulates the connection of the national power grid in the western United States. Relatively speaking, the distribution of nodes and edges is relatively balanced, which is likely to cause the infection process to end prematurely, and the node infection time is not very different. Therefore, this also leads to the loss based on time comparison to have a relatively lower influence on the overall loss.

[0100] Figure 7 Given the influence of different proportions of observation nodes and the relationship between the average length of random walk generated paths on the average error distance when the source is 2 on different networks (BA, WS, Email-Eu-core, ego Facebook, USPower Grid, Wiki-Vote). It can be seen that when the number of observation nodes is small and the detection degree of its neighborhood is low, the accuracy of the algorithm is basically beyond 3 hops, but as the detection range increases, that is, the average path length formed by each node increases, the accuracy of the algorithm is significantly improved. This phenomenon is more obvious in real networks such as Email and Facebook. This is probably because their connectivity is better than that of virtual networks, with more edges, and higher average path length and average clustering coefficient. From a vertical perspective, when the average path length detected reaches more than 20, the accuracy of the algorithm is significantly improved, and the average error distance drops to about 2 hops, and after that, the improvement is relatively small. At the same time, for the Wiki-Vote network with a large number of nodes and edges, the DL-SIA algorithm can also be greatly improved after the average path length detected reaches 30. In general, this method has a low detection cost and minimizes the amount of known information and calculation while ensuring a certain accuracy. From a horizontal perspective, the impact of the increase in observation nodes on the DL-SIA algorithm is weaker than the step size, but it still improves. When the proportion of observation nodes reaches about 30%-40%, the average error distance of the algorithm has a significant decrease. Increasing the proportion of observation nodes thereafter can also have a positive impact on the accuracy of the algorithm, but the benefits are relatively reduced.

[0101] Therefore, considering the practical factors and economic cost, it can be seen that the optimal ratio of DL-SIA observation nodes is about 40%-60%, and the detection radius is about 30-50. The above parameters can be adjusted appropriately according to the size of the network.

[0102] Figure 8 Given different algorithms, the running time comparison results in different networks are given. Figure 8 It can be seen that although the accuracy of the DMP algorithm is better than that of EPA and other algorithms, it takes twice as much time as other algorithms. This is related to the fact that it performs state reasoning for each node at any time t and does not adopt any pruning strategy. Centrality only considers a certain attribute of the node and performs unified calculations, which leads to its relatively superior running time, but its accuracy and average error distance are relatively poor. It can be said that DL-SIA still maintains a high accuracy rate for various situations while keeping the running time relatively short.

[0103] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present invention (including the claims) is limited to these examples. Under the concept of the present invention, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the present invention as described above, which are not provided in detail for the sake of simplicity.

[0104] The present invention is intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multi-source localization method based on deep learning, It is characterized in that The following steps are involved: Randomly select k propagation source nodes s* in the original undirected network; Based on the epidemic propagation SI model, the original network is diffused until the round ends or no new infected nodes are generated. Then the nodes and edges generated after the diffusion of the source nodes are extracted to generate the infected subgraph G infected ; According to the infected subgraph, a set proportion of infected nodes are randomly selected as the observer node set O = {o 1 , o 2 , ..., o m }, record the infection time of the selected node and obtain the infection time set T; Construct node feature extraction: Specifically, first consider the node propagation path, and then extract the observed infected nodes o i Simulate the infection path and obtain k paths l of set length through random walk, expressed as Generate path set C i , expressed as Extract C i The parameter α of all nodes in the graph relative to the observation node is denoted as According to the set C i and the set O computes the set of nodes V = {v 1 , v 2 , …, v k The characteristic vector group x of each node v in v , expressed as The calculation formula conforms to: where t i is the infection time corresponding to the i-th node in the infection time set T; Constructing a self-encoding network framework for v i The original feature values ​​are continuously compiled, integrated, and decoded, and the loss of each round of iteration is calculated to generate the integrated feature z i ; Among them, from four aspects, encoding loss, α difference loss, time difference loss and regularization term, the loss function of the self-encoding network is defined to generate the final integrated feature z i ; Combined with the generated integrated features z i Perform probability calculation to obtain the likelihood evaluation score of the node as the source, and select the first k nodes as the estimated source set.

2. The multi-source positioning method based on deep learning according to claim 1, It is characterized in that The self-encoding network framework is constructed for v i The original feature values ​​are continuously compiled, integrated, and decoded, and the loss of each round of iteration is calculated to generate the final integrated feature z i include: After the node feature extraction, the overall information of all nodes enters the multi-layer autoencoder network for training, and the encoding, integration, and decoding processes are repeated until the number of iterations is met or the loss reaches the threshold; The overall autoencoder network consists of two parts: the encoder and the decoder. The two layers are connected by the node integration features required by the positioning algorithm, that is, the integration layer. At the same time, the loss function of the autoencoder network is defined from four aspects: encoding loss, α difference loss, time difference loss and regularization term, to generate the final integrated feature z i .

3. The multi-source positioning method based on deep learning according to claim 2, It is characterized in that The method further comprises: Constructing a self-encoding network framework for v i The original feature value is compiled, and the encoding layer is To iterate, the decoding layer iterates with the following formula group: Calculate the loss function obtained in each round of iteration to control the number of iterations. The loss function consists of four parts: Loss of the x encoding x ; Code z n With z v The difference between uv Matching loss Loss α ; Node v and its C i Medium i Time matching loss o ; Regularization term L reg , calculated by the offset between each iteration and the previous iteration; the final loss is the sum of the four losses, where w and b are weights and offsets, respectively, which are randomly generated at the beginning of the neural network and gradually determined as the training proceeds; yv is the transfer vector of the v node vector in the hidden layer of the encoding layer, is the transfer vector in the Lth hidden layer of the coding layer.

4. The multi-source positioning method based on deep learning according to claim 3, It is characterized in that The calculation formula of comprehensive loss Loss is: Loss=Loss x +αLoss α +βLoss o +γLoss reg Among them, α, β and γ are hyperparameters preset by experiments, z is the output result of the network integration layer, and m is the number of hidden layers of the neural network.

5. The multi-source positioning method based on deep learning according to claim 4, It is characterized in that Generate integrated feature z i The formula is:

6. The multi-source positioning method based on deep learning according to claim 5, It is characterized in that Combined with the generated integrated features z i The formula for probability calculation is:

7. The multi-source positioning method based on deep learning according to claim 1, It is characterized in that The diffusion of the original network based on the epidemic transmission SI model specifically includes: Change the status of all nodes in the set to I state, and record the infection time as t 0 , enter the diffusion process; Set the threshold probability p that a node can be successfully infected infected In each round, the node marked as I state propagates to all its neighbors with its own probability, that is, the edge weight. If the threshold is met, the propagation is successful, and the state of the infected node changes from S state to I state. The above process continues until the end of the round or no new nodes are infected; After the infection process is completed, the original graph is pruned to extract the infected nodes and related edges to obtain the infected subgraph G infected .

Citation Information

Patent Citations

  • Social network information spreading source solving method based on random walk

    CN106557985A

  • Internet rumor detection method based on transmission graph neural network

    CN112732906A