A method for suppressing the spread of malicious entities in a social network
By applying SIR model and deep learning technology to calculate the weight of the edge connection, the prediction and suppression of bad entity propagation in social networks is solved, the effect of effectively suppressing bad entity propagation is achieved, and the trust in information propagation is improved.
Patent Information
- Application Number
- CN202211041720.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-29
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-08-29
AI Technical Summary
The prior art is difficult to effectively predict and suppress the process of spreading bad entities (such as computer viruses, malicious rumors) in social networks through the joint border, resulting in a decrease in trust in information transmission and potential social and economic damage.
The classic SIR infectious disease model is used to simulate the transmission process of bad entities, combined with the deepwalk algorithm and the natural language processing word2vec model, calculate the similarity of connected edges in the network as its weight value, and sort and delete key connected edges to suppress the transmission of bad entities.
Without significantly changing the network topology, effectively suppress the spread of bad entities, increase the trust in information dissemination, and reduce social and economic damage.
Smart Images

Figure CN115426153B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data mining and network edge processing technologies, and in particular to a method for suppressing the spread of malicious entities in a social network. Background Art
[0002] In real life, many physical scenarios can be expressed in the form of a network, such as computer networks, social networks, and the World Wide Web. The research on the structure and function of complex networks is also increasing. From a functional perspective, social network platforms play a key role in large-scale information dissemination. For example, the dissemination of positive content such as hot topics, opinions, and knowledge on online social networks enriches people's entertainment life. However, the dissemination of information in social networks is also accompanied by the generation of a large number of malicious entities. A malicious entity refers to something in the network that is likely to cause negative impacts. For example, computer viruses can spread through computer networks and email networks, and malicious rumors can spread among individuals through social networks. In these processes, the objects of dissemination are considered harmful and thus are not welcome.
[0003] On social networks, some users may deliberately spread rumors to deceive other users, or spread rumors inadvertently due to misunderstandings. These intentional or unintentional rumor spreads will lead to a decrease in users' trust in the news and messages disseminated in social networks, and may have a considerable impact on the credibility of the country and the government. It may even cause significant damage and economic losses, which is not conducive to social stability and citizen safety. Therefore, developing effective strategies to prevent malicious entities from spreading through the network is an important research issue.
[0004] Existing work reduces the spread range by removing nodes from the network. In particular, relevant research has shown that the strategy of removing nodes in descending order of out-degree can usually block or delay the spread of malicious entities. The process of deleting nodes is accompanied by deleting edges. That is to say, the task of deleting edges is more fundamental than the task of deleting nodes. Therefore, how to prevent the spread of malicious entities by removing edges from the underlying network is an important research topic. Summary of the Invention
[0005] The present invention aims to overcome the above-mentioned drawbacks of the prior art and proposes a method for suppressing the spread of malicious entities in a social network.
[0006] The present invention uses the classic SIR epidemic model to simulate the process of healthy nodes in the network being infected by the spread of malicious entities, records the spread paths of malicious entities and the IDs of infected nodes, and combines the deepwalk algorithm and the natural language processing word2vec model to calculate the similarity of edges in the network as their weight values, effectively solving the problem that the importance of edges cannot be effectively predicted due to insufficient network features.
[0007] The technical solution adopted by the present invention to achieve the above-mentioned invention purpose is as follows:
[0008] A method for suppressing the spread of malicious entities in a social network, comprising the following steps:
[0009] S1: Assume there is an undirected unweighted network G. Based on the random walk strategy of the SIR model, assume that a certain node in network G is the starting point of the spread of malicious entities. Conduct random walk sampling in the graph, record the spread path of malicious entities starting from this point, construct a corpus, and at the same time record the list of node IDs infected by this starting point. Repeat the above operations until all nodes are traversed;
[0010] S2: Place the corpus described in step S1 into the word2vec model for training. Use the trained model to construct a node feature matrix, calculate the node similarity matrix, and obtain the similarity sim1 of each pair of nodes connected by edges in network G;
[0011] S3: According to the list of infected node IDs described in step S1, calculate the number of the same infected node IDs existing between the nodes connected by each edge in network G, and calculate the similarity sim2 of each pair of nodes connected by edges in network G;
[0012] S4: Take the weighted average of sim1 and sim2 as the weight value of the edge in network G. Sort the edge weight values, and deleting the top-k edges can effectively suppress the spread of malicious entities.
[0013] Preferably, in step S1: Assume there is an undirected unweighted social network G=(V, E), where V is the set of nodes in the network and E is the set of edges in the network. Construct the SIR model, set the time window to days = 100, set the infection rate β = 0.1, set the recovery rate γ = 0.05. Take a certain node in the network as the initial infected node, and the remaining nodes as susceptible nodes. Record the infection spread path of this node, construct a corpus, and at the same time record the list of node IDs infected by this starting point. Repeat the above operations until all nodes in network G are traversed.
[0014] Preferably, in step S2: Organize all the above-mentioned spread paths into a list format, use it as a corpus to input into the word2vec model, set vector_size = 64, that is, the dimension of node embedding is 64. Construct a vocabulary and train, save the trained model, construct a node feature matrix, whose dimension should be |V|×64, and calculate the cosine similarity between nodes as shown in the following formula (1):
[0015]
[0016] Among them, X and Y are 64-dimensional node feature vectors respectively. A symmetric matrix |V|×|V| with cosine similarity as matrix elements is constructed. The value corresponding to the a-th row and the b-th column in the similarity matrix can be retrieved according to the edge [a, b] in the network G, denoted as sim1 a,b , which is a similarity of the edge [a, b].
[0017] Preferably, in the step S3: according to the list of infected node IDs starting from different nodes in the network G described in the step S1, calculate the number of the same infected node IDs between the node pairs of each edge in the network G, denoted as num, and construct a nested list of node pairs and the number of the same node IDs, such as [[a, b, numa, b],...], where a and b are nodes, and num a,b represents the number of the same infected node IDs between the node pair [a, b], and the similarity sim2 of the node pair [a, b] in the network G is calculated by normalization a,b , as shown in the following formula (2):
[0018]
[0019] where num a,b represents the number of the same infected node IDs between the node pair [a, b], and num max and num min represent the maximum and minimum values of the number of the same infected nodes among all node pairs.
[0020] Preferably, in the step S4: use the weighted average of sim1 a,b and sim2 a,b as the weight value of the edge in the network G, as shown in the following formula (3):
[0021] sim a,b = 0.5×sim1 a,b + 0.5×sim2 a,b (3)
[0022] Use sim a,b as the weight value of the edge and write it into the edge list of the network G. Sort all edges according to the weight value from large to small, and deleting the top-k edges can effectively inhibit the spread of malicious entities.
[0023] With the rise of Internet media, the security problems brought by the spread of malicious entities in social networks have become increasingly prominent. In order to inhibit the wanton spread of malicious entities in the network space on the premise of changing the network topology as little as possible, the present invention effectively utilizes the topological features of the network, predicts the key edges for the spread of malicious entities in the network, and blocking the key edges can effectively inhibit the spread of malicious entities.
[0024] The advantages of the present invention are as follows: on the premise of fewer network features, by using the topological structure of the network, it can ensure that the spread of malicious entities is effectively suppressed only under minor perturbations. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flowchart of the method of the present invention;
[0026] Figure 2 Performing the edge deletion operation on the Wikipedia dataset, the effect of blocking 100 edges using the method of the present invention is the same as the effect of blocking more than 1800 edges using the random method. DETAILED DESCRIPTION OF THE INVENTION
[0027] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0028] This embodiment provides a method for suppressing the spread of false messages in a social network, which is based on the method for suppressing the spread of malicious entities in a social network according to the present invention.
[0029] In order to efficiently suppress the spread of malicious entities in the social network, the present invention uses the classical SIR epidemic model to simulate the process of healthy nodes in the network being infected by the spread of malicious entities, records the spread path of malicious entities and the IDs of the infected nodes, and combines the deepwalk algorithm and the natural language processing word2vec model method to calculate the similarity of the edges in the network as its weight value, and the proposed method effectively solves the problem that it is impossible to effectively predict the importance of edges due to insufficient network features.
[0030] The technical solution adopted by the present invention to achieve the above object is as follows:
[0031] A method for suppressing the spread of malicious entities in a social network, comprising the following steps:
[0032] S1: Assume that there is an undirected unweighted network G. Based on the random walk strategy of the SIR model, assume that a certain node in the network G is the starting point of the spread of malicious entities, perform random walk sampling in the graph, record the spread path of malicious entities starting from this point, construct a corpus, and at the same time record the list of node IDs infected by this starting point. Repeat the above operations until all nodes are traversed;
[0033] S2: Place the corpus described in step S1 in the word2vec model for training. Using the trained model, construct a node feature matrix, calculate the node similarity matrix, and obtain the node pair similarity sim1 of each edge in the network G;
[0034] S3: According to the list of infected node IDs starting from different nodes in the network G described in step S1, calculate the number of the same infected node IDs existing between the node pairs of each edge in the network G, and calculate the similarity sim2 of the node pairs of each edge in the network G after normalization;
[0035] S4: Take the weighted average of sim1 and sim2 as the weight value of the edge in the network G, sort the edge weight values, and deleting the top-k edges can effectively inhibit the spread of malicious entities.
[0036] In step S1: Assume there is an undirected unweighted social network G = (V, E), where V is the set of nodes in the network and E is the set of edges in the network. Construct an SIR model, set the time window as days = 100, set the infection rate β = 0.1, set the recovery rate γ = 0.05. Take a certain node in the network as the initial infected node, and the remaining nodes as susceptible nodes. Record the infection propagation path of this node, construct a corpus, and at the same time record the list of node IDs infected by this starting point. Repeat the above operations until all nodes in the network G are traversed.
[0037] In step S2: Organize all the above-mentioned propagation paths into a list format, input it into the word2vec model as a corpus, set vector_size = 64, that is, the dimension of node embedding is 64, construct a vocabulary and train it, save the trained model, construct a node feature matrix, whose dimension should be |V|×64, and calculate the cosine similarity between nodes, as shown in the following formula (1):
[0038]
[0039] where X and Y are 64-dimensional node feature vectors respectively. Construct a symmetric matrix |V|×|V| with cosine similarity as matrix elements, and the value corresponding to the a-th row and b-th column in the similarity matrix can be retrieved according to the edge [a, b] in the network G as sim1 a,b , which is a similarity of the edge [a, b].
[0040] In step S3: According to the list of infected node IDs starting from different nodes in the network G described in step S1, calculate the number of the same infected node IDs existing between the node pairs of each edge in the network G, denoted as num, and construct a nested list of node pairs - the number of the same node IDs, such as [[a, b, num a,b ,...], where a and b are nodes, and num a,b represents the number of the same infected node IDs between the node pair [a, b], and calculate the similarity sim2 of the node pairs [a, b] of each edge in the network G after normalization a,b , as shown in the following formula (2):
[0041]
[0042] Among them, num a,b represents the number of identical infected node IDs between node pairs [a, b], num max and num min represent the maximum and minimum values of the number of identical infected nodes among all node pairs.
[0043] In step S4: Take the weighted average of sim1 a,b and sim2 a,b as the weight value of the edges in network G, as shown in the following formula (3):
[0044] sim a,b = 0.5 × sim1 a,b + 0.5 × sim2 a,b (3)
[0045] Take sim a,b as the weight value of the edges and write it into the edge list of network G. Sort all the edges according to the weight value from large to small, and deleting the top-k edges can effectively inhibit the spread of malicious entities.
[0046] A method for suppressing the spread of malicious entities in a social network according to the present invention samples network topology information through the random walk strategy of the SIR model, constructs a similarity matrix of network nodes by referring to the method of natural language processing, and calculates the node similarity. The final result shows that a method for suppressing the spread of malicious entities in a social network has excellent effects, can successfully inhibit the spread of false messages in the social network, and the final result based on this method can play an important role in the information dissemination direction of the social network. Therefore, effectively suppressing the spread of false messages on the network is particularly important in the field of network security.
[0047] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments, and the protection scope of the present invention also extends to equivalent technical means that those skilled in the art can think of according to the inventive concept of the present invention.
Claims
1. A method for suppressing the spread of malicious entities in a social network, characterized in that: It includes the following steps: S1: Construct an undirected unweighted network G. Based on the random walk strategy of the SIR model, using a certain node in network G as the starting point of the spread of malicious entities, conduct random walk sampling in the graph, record the spread path of malicious entities starting from this point, construct a corpus, and at the same time record the list of node IDs infected by this starting point. Repeat the above operations until all nodes are traversed; S2: Place the corpus described in step S1 into the word2vec model for training. Use the trained model to construct a node feature matrix, calculate the node similarity matrix, and obtain the similarity sim1 of node pairs for each edge in network G; Specifically, it includes: Form a list of all the above-mentioned spread paths as the corpus and input it into the word2vec model. Set vector_size = 64, that is, the dimension of node embedding is 64. Construct a vocabulary and train, save the trained model, construct a node feature matrix, whose dimension should be |V|×64, where V is the set of nodes in the network. Calculate the cosine similarity between nodes as shown in the following formula (1): Where X and Y are 64-dimensional node feature vectors respectively, a similarity matrix |V|×|V| with cosine similarity as matrix elements is constructed, and according to the edge [a, b] in the network G, the value corresponding to the a-th row and the b-th column in the similarity matrix is retrieved as sim1 a,b , which is a similarity of the edge [a, b]; S3: According to the list of node IDs infected starting from different nodes in network G described in step S1, calculate the number of node pairs with the same infected node ID for each edge in network G, and calculate the similarity sim2 of node pairs for each edge in network G through normalization; Specifically, it includes: According to the list of infected node IDs starting from different nodes in the network G described in step S1, calculate the number of pairs of nodes with the same infected node ID among the node pairs of each edge in the network G, denoted as num, and construct a nested list [[a, b, num a,b , …], where a and b are nodes, and num a,b represents the number of the same infected node IDs between the node pair [a, b], and calculate the similarity sim2 of each edge node pair [a, b] in the network G by normalization a,b , as shown in the following formula (2): Among them, num a,b represents the number of nodes with the same infected node ID between node pairs [a, b], num max and num min represent the maximum and minimum values of the number of nodes with the same infection among all node pairs; S4: Take the weighted average of sim1 and sim2 as the weight value of the edges in network G, sort the edge weight values, and deleting the top-k edges can effectively inhibit the spread of malicious entities.
2. The method for suppressing the spread of malicious entities in a social network according to claim 1, characterized in that: The step S1 includes: Construct an undirected unweighted social network G=(V, E), where E is the set of edges in the network. Construct the SIR model, set the time window as days = 100, set the infection rate β = 0.1, set the recovery rate γ = 0.
05. Use a certain node in the network as the initial infected node and the remaining nodes as susceptible nodes. Record the infection spread path of this node, construct a corpus, and at the same time record the list of node IDs infected by this starting point. Repeat the above operations until all nodes in network G are traversed.
3. The method for suppressing the spread of malicious entities in a social network according to claim 1, characterized in that: The step S4 includes: Take the weighted average of sim1 a,b and sim2 a,b as the weight of the edge in network G, as shown in the following formula (3): sim a,b = 0.5 × sim1 a,b + 0.5 × sim2 a,b (3) Take the sim a,b as the weight of the connected edge and write it into the edge list of network G. Sort all connected edges according to the weight from large to small. Deleting the top-k connected edges can effectively suppress the spread of malicious entities.
Citation Information
Patent Citations
Social network credibility learning method based on weight update
CN108334953A
K-shell decomposition-based method and K-shell decomposition-based device for identifying propagation key node
CN110247805A