A worm propagation tracing method based on a convolutional neural network model

Through the method based on convolutional neural network, the implicit propagation prior knowledge is learned from the propagation graph samples, and the problem of dependence on the propagation model prior knowledge in the existing technology is solved, achieving a more accurate worm propagation traceability effect.

CN115484079BActive Publication Date: 2025-05-30SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211059770.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-05-30
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

Existing worm transmission traceability methods rely on prior knowledge of the dissemination model, making it difficult to obtain accurate traceability results under complex and changeable worm transmission.

Method used

The method based on convolutional neural network is adopted to learn implicit propagation prior knowledge from the propagation graph samples of a large number of known propagation sources, establish a corresponding relationship model between the propagation graph structural characteristics and the propagation source nodes, and convert it into a supervised propagation graph classification problem to be solved.

Benefits of technology

It has got rid of the dependence on the prior hypothesis of the propagation model, and has improved the accuracy and effectiveness of worm transmission traceability, and is suitable for worm transmission traceability under different observation conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115484079B_ABST
    Figure CN115484079B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for tracing the source of worm propagation based on a convolutional neural network model, including: 1. Collecting a sample set of worm propagation graphs, simulating the worm propagation process using the SI model, and obtaining a sample set of worm propagation graphs with different nodes as source nodes; 2. Respectively under the conditions of complete observation, snapshot observation, and sensor observation, converting the propagation graph samples in the non-Euclidean space into a two-dimensional matrix in the Euclidean space in the form of an adjacency graph; 3. Using the converted two-dimensional matrix of the propagation graph as the input of the convolutional neural network model (CNN), and taking the source node corresponding to the propagation graph as the class label output of the graph, and training the CNN using the gradient descent algorithm based on the sample set of propagation graphs; 4. Inputting the propagation graph with an unknown propagation source into the trained convolutional neural network to obtain the prediction result of its propagation source node (i.e., the tracing result). This method solves the problem of tracing the source of Internet worm propagation from the perspective of supervised propagation graph classification based on the convolutional neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for tracing the spread of worms based on a convolutional neural network model, which is applicable to worms widely spread in the Internet, performs reverse tracking and tracing on the worm spread, curbs the worms in the initial stage of their spread, and provides key evidence support for subsequent virus feature analysis and Internet crime forensics. Background Art

[0002] With the continuous advancement of the informatization process in China, a large number of industrial enterprises and government agencies are committed to relying on Internet technology to achieve information sharing and interconnection of their management systems and key infrastructure. While these infrastructures are exposed to the Internet, they also bring network security problems. Especially by leveraging the connectivity of the Internet, worms can spread widely and then cause serious consequences. Traditional tracing methods simulate the spread process based on the prior knowledge of the assumed spread model, estimate the spread source node through probability reasoning and graph analysis. Only when the spread assumption conforms to the actual situation, the tracing method is effective. However, in actual situations, it is difficult to obtain the real spread model, which poses new challenges to worm spread tracing.

[0003] On the one hand, existing worm spread tracing methods often assume prior knowledge of the spread model and find the spread source node by reproducing the spread process. Such methods are effective for spreads where the assumed model conforms to the actual situation. However, in reality, there are many types of worms, and their spread situations are relatively complex and changeable, making it difficult to obtain the real spread model. In addition, the more assumption parameters for the spread process, the more accurate the description will be, and the higher the accuracy of the tracing method. Therefore, the prediction accuracy of traditional tracing methods is limited. Generally, the obtained estimation results are not the real source nodes but estimation nodes close to the real source nodes.

[0004] On the other hand, in recent years, convolutional neural networks can be widely applied to solve classification problems of Euclidean space information such as images, voices, and structured texts. However, the spread graph obtained from worm spread is data in a non-Euclidean space, with arbitrary sizes and complex topological structures. Therefore, when applying convolutional neural networks, the data in the non-Euclidean space needs to be converted into a two-dimensional matrix in the Euclidean space in the form of an adjacency graph. On this basis, the converted two-dimensional matrix is used as the input of the convolutional neural network, and the source node corresponding to the spread graph is used as the class label of the graph. Based on the gradient descent algorithm, the implicit prior knowledge of the spread model is learned from a large number of spread graph samples with known spread sources, and the unsupervised worm spread tracing problem is transformed into a supervised spread graph classification problem for solution.

[0005] In summary, to solve the problem of worm spread tracing for Internet security, it is necessary to break through the dependence on prior spread assumptions and learn the prior knowledge of the spread model based on convolutional neural networks. Summary of the Invention

[0006] The object of the present invention is to get rid of the dependence on the prior assumption of the propagation model in traditional traceability problems, and propose a worm propagation traceability method based on a convolutional neural network model. By learning the implicit propagation prior knowledge from a large number of propagation graph samples with known propagation sources, a correspondence model between the propagation graph structure features and the propagation source nodes (class labels) is established, so as to convert the unsupervised propagation traceability problem into a supervised propagation graph classification problem for solution.

[0007] In order to achieve the above object of the invention, the present invention is realized through the following specific technical solutions:

[0008] A worm propagation traceability method based on a convolutional neural network model includes the following steps:

[0009] Step 1) Collect a set of worm propagation graph samples. The present invention uses the SI (S represents the node in the susceptible state, I represents the node in the infected state) model in the infectious disease model to simulate the worm propagation process, and obtains propagation graph samples of different nodes as source nodes propagating on the network;

[0010] Step 2) Under the conditions of complete observation, snapshot observation and sensor observation respectively, convert the propagation graph samples in the non-Euclidean space into a two-dimensional matrix in the Euclidean space in the form of an adjacency graph;

[0011] Step 3) Use the converted two-dimensional matrix of the propagation graph as the model input of the convolutional neural network, use the corresponding source node of the propagation graph as the graph class label output, and train the neural network using the gradient descent algorithm;

[0012] Step 4) Input the propagation graph with an unknown propagation source into the neural network to obtain the prediction result of its propagation source node.

[0013] The specific steps of the said step 1) include the following steps:

[0014] Step 1.1: The propagation base map is a directed graph, and all edges are randomly assigned weights weight as the probability of being infected between the nodes of the edges;

[0015] Step 1.2: During the propagation process, the propagation infection probability is randomly set to q, which follows a uniform distribution on (0, 1). When q > weight, the node is infected. After a period of time, a worm propagation graph sample starting from the source node s can be obtained;

[0016] Step 1.3: In the actual propagation process, the specific infection time of the node cannot be obtained, but the infection scale of the node can be observed, that is, when the number of infected nodes reaches a certain range, the propagation stops.

[0017] The specific steps of the said step 2) include the following steps:

[0018] Step 2.1: Under the condition of complete observation, fix the node numbers of the propagation base map and observe the states of all nodes. Assign a value of 0 to the nodes in the susceptible state and a value of 1 to the nodes in the infected state.

[0019] Step 2.2: Obtain the weights of the edges connected between nodes according to the sum of node states. The edges connected between susceptible nodes are represented by 0, the edges connected between a susceptible node and an infected node are represented by 1, and the edges connected between two infected nodes are represented by 2.

[0020] Step 2.3: Under the condition of complete observation, represent the propagation graph as a two-dimensional matrix. For the propagation base map with N nodes, the size of the two-dimensional matrix is N×N.

[0021] Step 2.4: Under the conditions of snapshot observation and sensor observation, only the infection states of some nodes can be observed. When the number of observed nodes is M, the corresponding size of the two-dimensional matrix is M×M.

[0022] Step 2.5: Under the condition of sensor observation, the infection time of nodes can be observed, and the state of the edge is represented by the sum of the node infection times. Convert the propagation graph into a two-dimensional matrix.

[0023] The specific steps of step 3) are as follows:

[0024] Step 3.1: The convolutional neural network includes a convolutional layer and a fully connected layer. Combine the convolutional layer, activation function, and pooling layer into a group of stacked layers. Use two stacked layers and one fully connected layer as the CNN model.

[0025] Step 3.2: Select a convolutional kernel with a size of 3×3 in the convolutional layer. The number of convolutional kernels in the two convolutional layers is 16 and 32 respectively. The activation function is Relu(), and the filter size of the pooling layer is 2×2.

[0026] Step 3.3: The input of the neural network is a two-dimensional matrix, and the output is the source node label corresponding to the two-dimensional matrix. Use CNN to learn the prior knowledge of the propagation model from the propagation graph samples and establish a correspondence model between the propagation graph structure features and the propagation source nodes.

[0027] Step 3.4: During the training of the neural network, calculate the difference between the output label obtained through forward propagation and the actual source node label to obtain the loss of the neural network training.

[0028] Step 3.5: Backpropagate the network loss using the gradient descent method to update the weight values of the edges of the neural network. Repeat steps 3.3 - 3.5 until the network loss converges.

[0029] The specific steps of step 4) are as follows:

[0030] Step 4.1: For the propagation graph with unknown propagation sources, input it into the trained neural network to obtain the predicted source node labels.

[0031] Step 4.2: Calculate the shortest distance between the predicted source node and the actual source node as the error distance. The smaller the error distance, the better the prediction effect. For a large number of unknown propagation source samples, calculate the average error distance and prediction accuracy simultaneously to evaluate the prediction effect of the algorithm.

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] It is not limited to the dependence on the prior assumptions of the propagation model. By learning the implicit prior knowledge of the propagation model through a convolutional neural network, a corresponding relationship model between the propagation graph structure features and the propagation source nodes is established. It can learn from the propagation graph samples with known propagation sources, predict the source nodes for the propagation graph samples with unknown propagation sources, and at the same time, it is discussed that the worm propagation traceability method based on the CNN model is also applicable under different observation conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of the method of the present invention.

[0035] Figure 2 It is the overall flowchart of the method of the present invention.

[0036] Figure 3 It is a schematic diagram of converting a fully observed propagation graph into a two-dimensional matrix.

[0037] Figure 4 It is a schematic diagram of converting a partially observed (snapshot observation, sensor observation) propagation graph into a two-dimensional matrix.

[0038] Figure 5 It is a schematic diagram of the worm propagation traceability method based on the CNN model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The following provides a preferred embodiment of the present invention in conjunction with the drawings to detail the technical solution of the present invention.

[0040] The object of the present invention is to directly learn the prior knowledge of the propagation model from the propagation graph samples through a supervised traceability method based on the CNN model, and solve the worm propagation traceability problem from the perspective of graph classification, so as to have the opportunity to obtain a more accurate worm propagation prior model.

[0041] As Figure 1 and Figure 2 shown, this embodiment takes the propagation traceability of a 100-node BA scale-free network, a WS small-world network, and a 117-node Phy real social network as an example. The specific implementation steps are as follows:

[0042] Step 101: Take the BA100, WS100, and Phy117 networks as the propagation base maps respectively. Randomly assign weights weight to all edges as the probability of infection between the nodes of the edges, and represent the propagation base maps as weighted directed graphs;

[0043] Step 102: During the propagation process, randomly set the propagation infection probability to q, which follows a uniform distribution on (0, 1). When q > weight, the node is infected. After a period of time, a worm propagation graph starting from the source node s can be obtained. Randomly divide the data set into a training set of 50N and a test set of 5N, where N = 100, 100, 117 is the number of nodes, that is, 55 simulations are performed starting from each node to obtain the propagation graph;

[0044] Step 103: During the actual propagation process, the specific infection time of the nodes cannot be obtained, but the infection scale of the nodes can be observed. When the number of infected nodes is 30% of the total number of nodes, stop the propagation.

[0045] Step 201: Under the condition of complete observation, fix the node numbers of the propagation base map and observe the states of all nodes. Assign 0 to the nodes in the susceptible state and 1 to the nodes in the infected state;

[0046] Step 202: Obtain the weights of the edges connected between the nodes according to the sum of the node states. The edges connected between susceptible nodes are represented by 0, the edges connected between a susceptible node and an infected node are represented by 1, and the edges connected between two infected nodes are represented by 2;

[0047] Step 203: Represent the propagation graph as a two-dimensional matrix. For the propagation base map of N nodes, the size of the two-dimensional matrix is N×N, as Figure 3 shown;

[0048] Step 204: Under the conditions of snapshot observation and sensor observation, only the infection states of some nodes can be observed. When the number of observed nodes is M = 50%×N, the corresponding size of the two-dimensional matrix is M×M, as Figure 4 shown;

[0049] Step 205: Under the condition of sensor observation, the infection time of the nodes can be observed, and the state of the edges is represented by the sum of the node infection times, and the propagation graph is converted into a two-dimensional matrix.

[0050] Step 301: The convolutional neural network includes a convolutional layer and a fully connected layer. Combine the convolutional layer, activation function, and pooling layer into a group of stacked layers. Use two stacked layers and one fully connected layer as the CNN model, as Figure 5 shown;

[0051] Step 302: Select a convolution kernel of size 3×3 in the convolutional layer. The number of convolution kernels in the two convolutional layers is 16 and 32 respectively, the activation function is Relu(), and the filter size of the pooling layer is 2×2;

[0052] Step 303: The input of the neural network is a two-dimensional matrix N×N, and the output is the source node label corresponding to the two-dimensional matrix, with a size of N. Use CNN to learn the prior knowledge of the propagation model from the propagation graph samples, and establish a correspondence model between the propagation graph structure features and the propagation source nodes;

[0053] Step 304: When training the neural network, use cross-entropy to calculate the loss. That is, for a certain propagation graph, the output label O obtained from its input forward propagation t and the actual source node label O r The difference between them is quantified by cross-entropy, and the calculation formula is where j represents the jth propagation graph sample;

[0054] Step 305: Use the gradient descent method to update the model weights, that is where the model learning rate η = 0.001. Repeat steps 303 - 305 until the network loss converges. L represents the cross-entropy L(O r , O t ), and w represents the neural network model weights.

[0055] Step 401: For the propagation graphs of unknown propagation sources corresponding to different propagation base maps, input them into the corresponding trained neural network to obtain the predicted source node label O t ;

[0056] Step 402: Calculate the shortest distance between the predicted source node O t and the actual source node O r as the error distance. The smaller the error distance, the better the prediction effect. For a large number of unknown propagation source samples, calculate the average error distance (AverageError Distance, AED) and the prediction accuracy (Accuracy, ACC) at the same time to evaluate the prediction effect of the traceability algorithm.

[0057] The following table lists the experimental results of the method proposed in the present invention under the conditions of complete observation, snapshot observation, and sensor observation in the BA network, WS network, and the real social network Phy. The results show the effectiveness of the method of the present invention:

[0058]

[0059] The above-described embodiment is a worm propagation traceability method based on a convolutional neural network model. It collects a sample set of worm propagation graphs, uses the SI model to simulate the worm propagation process, and obtains a sample set of worm propagation graphs with different nodes as source nodes. Under the conditions of complete observation, snapshot observation, and sensor observation respectively, the propagation graph samples in the non-Euclidean space are converted into a two-dimensional matrix in the Euclidean space in the form of an adjacency graph. Then, the two-dimensional matrix after the conversion of the propagation graph is used as the input of the convolutional neural network model (CNN), and the source node corresponding to the propagation graph is output as the class label of the graph. Based on the sample set of propagation graphs, the CNN is trained using the gradient descent algorithm. Then, the propagation graph with an unknown propagation source is input into the trained convolutional neural network to obtain the prediction result of its propagation source node (i.e., the traceability result). The method of the above-described embodiment is based on the convolutional neural network model and solves the problem of Internet worm propagation traceability from the perspective of supervised propagation graph classification.

[0060] The specific embodiments described above further elaborate on the technical problems solved, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for tracing the source of worm propagation based on a convolutional neural network model, characterized in that, it includes the following steps: Step 1) Collect a sample set of worm propagation graphs, use the SI model to simulate the worm propagation process, and obtain a sample set of propagation graphs where different nodes act as source nodes and propagate on the network; Step 2) Under the conditions of complete observation, snapshot observation, and sensor observation respectively, convert the propagation graph samples in the non-Euclidean space into a two-dimensional matrix in the Euclidean space in the form of an adjacency graph; Step 3) Use the two-dimensional matrix after the propagation graph conversion as the model input of the convolutional neural network, and use the source node corresponding to the propagation graph as the class label output of the graph, and train the neural network using the gradient descent algorithm; Step 4) Input the propagation graph with unknown propagation source into the neural network to obtain the prediction result of its propagation source node.

2. The method for tracing the source of worm propagation based on a convolutional neural network model according to claim 1, characterized in that, the specific steps of the said Step 1) include the following steps: Step 1.1: The propagation base map is a directed graph, and all edges are randomly assigned weights weight as the probability of being infected between the nodes of the edges; Step 1.2: During the propagation process, randomly set the propagation infection probability to q, which follows a uniform distribution on (0,1). When q > weight, the node is infected. After a period of time, a worm propagation graph starting from the source node s can be obtained; Step 1.3: In the actual propagation process, the specific infection time of the nodes cannot be obtained, but the infection scale of the nodes can be observed, that is, when the number of infected nodes reaches a certain range, the propagation stops.

3. The method for tracing the source of worm propagation based on a convolutional neural network model according to claim 1, characterized in that, the specific steps of the said Step 2) include the following steps: Step 2.1: Under the condition of complete observation, fix the node numbers of the propagation base map and observe the states of all nodes. Assign 0 to the nodes in the susceptible state and 1 to the nodes in the infected state; Step 2.2: Obtain the weights of the edges connected between the nodes according to the sum of the node states. The edges connected between susceptible nodes are represented by 0, the edges connected between a susceptible node and an infected node are represented by 1, and the edges connected between two infected nodes are represented by 2; Step 2.3: Under the condition of complete observation, represent the propagation graph with a two-dimensional matrix. For a propagation base map with N nodes, the size of the two-dimensional matrix is N×N; Step 2.4: Under the conditions of snapshot observation and sensor observation, only the infection states of some nodes can be observed. When the number of observed nodes is M, the corresponding size of the two-dimensional matrix is M×M; Step 2.5: Under the condition of sensor observation, the infection time of the nodes can be observed, and the state of the edges is represented by the sum of the node infection times, and the propagation graph is converted into a two-dimensional matrix.

4. The method for tracing the source of worm propagation based on a convolutional neural network model according to claim 1, characterized in that, the specific steps of the said Step 3) include the following steps: Step 3.1: The convolutional neural network includes a convolutional layer and a fully connected layer. Combine the convolutional layer, activation function, and pooling layer into a group of stacked layers, and use two stacked layers and one fully connected layer as the CNN model; Step 3.2: Select a convolution kernel of size 3×3 in the convolutional layer. The number of convolution kernels in the two convolutional layers is 16 and 32 respectively. The activation function is Relu(), and the filter size of the pooling layer is 2×2; Step 3.3: The input of the neural network is a two-dimensional matrix, and the output is the source node label corresponding to the two-dimensional matrix. Use CNN to learn the prior knowledge of the propagation model from the propagation graph samples, and establish a correspondence model between the propagation graph structure features and the propagation source nodes; Step 3.4: When the neural network is trained, calculate the difference between the output label obtained through forward propagation and the actual source node label to obtain the loss of the neural network training; Step 3.5: Backpropagate the network loss using the gradient descent method to update the weight values of the edges of the neural network. Repeat steps 3.3 - 3.5 until the network loss converges.

5. The worm propagation traceability method based on the convolutional neural network model according to claim 1, wherein, the specific steps of step 4) are as follows: Step 4.1: For the propagation graph with unknown propagation source, input it into the trained neural network to obtain the predicted source node label; Step 4.2: Calculate the shortest distance between the predicted source node and the actual source node as the error distance. The smaller the error distance, the better the prediction effect. For a large number of unknown propagation source samples, calculate the average error distance and the prediction accuracy at the same time to evaluate the prediction effect of the algorithm.

Citation Information

Patent Citations

  • Communication source positioning method based on biological intelligence

    CN106355091A

  • Network virus tracing method and system, equipment, medium and processing terminal

    CN113114657A