Alarm Correlation Analysis Method Based on Graph Neural Network
Through the alarm correlation analysis method based on graph neural network, an attack graph is constructed and a classification model is trained, which solves the problem that the attack scenario cannot be accurately identified in the existing technology, and achieves more efficient attack scenario recognition and lower false alarm rate.
Patent Information
- Application Number
- CN202210835786.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-15
AI Technical Summary
The existing alarm association analysis methods cannot accurately identify attack scenarios, ignore the correlation between alarms, resulting in inefficient recognition.
The alarm correlation analysis method based on graph neural network is adopted, and the alarm data is preprocessed through the causal correlation module to construct an attack graph, and the image neural network module is used to train the graph neural network classification model to identify attack scenarios.
It improves the efficiency of alarm correlation analysis and model inference ability, improves the accuracy and recall rate of attack scene recognition, and reduces false positives and missed reports.
Smart Images

Figure CN115643153B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph neural networks, and in particular to an alarm correlation analysis method based on graph neural networks. Background Art
[0002] In recent years, the network system structure and intrusion means have become increasingly diverse and complex, and the analysis cost of IDS for alarm logs has also increased exponentially. According to research, 84% of attackers will leave security evidence when implementing destruction, but the detected alarms are often only a certain activity in the attack chain. Therefore, correlating and analyzing security events can manage network systems more effectively. However, past attack scenario recognition methods have ignored the correlation between alarms.
[0003] Currently, the research on alarm correlation analysis is not yet perfect. Traditional methods are objective and have reasoning ability, but the correlation results are not satisfactory and often cannot accurately identify the attack scenarios to which the alarms belong. Summary of the Invention
[0004] The purpose of the present invention is to provide an alarm correlation analysis method based on graph neural networks, aiming to solve the problem that existing analysis methods cannot accurately identify attack scenarios.
[0005] To achieve the above purpose, the present invention provides an alarm correlation analysis method based on graph neural networks, including the following steps:
[0006] Preprocess the alarm data through a causal correlation module to obtain an attack graph;
[0007] The image neural network module extracts the attack graph to train a graph neural network, obtaining a graph neural network classification model;
[0008] Identify test data through the graph neural network classification model to obtain an attack scenario.
[0009] Among them, the specific method of preprocessing the alarm data through the causal correlation module to obtain an attack graph:
[0010] The causal correlation module performs standardized cleaning on the alarm data to obtain processed data;
[0011] Construct the attack graph based on the processed data using a causal correlation analysis method.
[0012] Among them, the specific method of the image neural network module extracting the attack graph to train a graph neural network and obtaining a graph neural network classification model:
[0013] The image neural network module abstracts the attack graph into an adjacency matrix;
[0014] Build the initial network structure of the graph neural network based on the adjacency matrix;
[0015] Introduce the SSA parameter into the initial network structure to train the initial network structure, and obtain the graph neural network classification model.
[0016] Among them, the visualization processing software is the drawing software graphviz.
[0017] Among them, the specific method of introducing the SSA parameter into the initial network structure to obtain the graph neural network classification model:
[0018] Set the model Loss as the fitness function of SSA, and specify the parameter optimization range to obtain the SSA parameter;
[0019] Introduce the SSA parameter into the initial network structure, and train the initial network structure to obtain the graph neural network classification model.
[0020] The alarm correlation analysis method based on the graph neural network of the present invention preprocesses alarm data through a causal correlation module to obtain an attack graph; the image neural network module extracts the information of the attack graph, trains the graph neural network, and obtains a graph neural network classification model; identifies data through the graph neural network classification model to obtain an attack scenario. This method first analyzes the attack scenario, designs the matching rules for the preconditions and results of security events; then uses the causal correlation analysis method to obtain related alarm sequences; visualizes the attack graph using a drawing tool, prepares the input data of the graph neural network, and extracts the information of the attack graph; builds the initial network structure of the graph neural network and trains the graph neural network classification model; finally, identifies the attack scenario to which the test alarm belongs. This method further extracts the information of the attack graph, improves the efficiency of attack scenario recognition, ensures the correlation efficiency while improving the inference ability of the model, and solves the problem that the existing analysis methods cannot accurately identify the attack scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 It is a schematic diagram comparing the classification performance of different classifiers.
[0023] Figure 2 It is a schematic diagram of the Loss convergence curve.
[0024] Figure 3 It is a schematic diagram of the performance comparison of the alarm correlation analysis method based on different graph neural networks.
[0025] Figure 4 It is a schematic diagram of the SSA-GCN structure.
[0026] Figure 5 It is a schematic diagram of the performance comparison of the alarm correlation analysis method based on different graph neural networks.
[0027] Figure 6 It is an architecture diagram of intelligent alarm correlation analysis based on graph neural networks.
[0028] Figure 7 It is a process diagram of alarm data correlation analysis.
[0029] Figure 8 It is a flowchart of the alarm correlation analysis method based on graph neural networks provided by the present invention. Detailed implementation manners
[0030] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0031] Please refer to Figures 1 to 8 , the present invention provides an alarm correlation analysis method based on graph neural networks, including the following steps:
[0032] S1 Preprocess the alarm data through a causal association module to obtain an attack graph;
[0033] Specifically, first analyze the attack scenario and design matching rules for the preconditions and results of security events; then use the causal association analysis method to obtain related alarm sequences; finally, use the drawing tool graphviz to visualize the network attack graph. The visualization processing software is the drawing software graphviz, clean the alarm data such as standardization, and generate an alarm attack graph and an attack scenario using the causal association analysis method at this stage. Otherwise, the graph neural network (GNN) cannot be trained.
[0034] Specific manner:
[0035] S11 The causal association module performs standardized cleaning on the alarm data to obtain processed data;
[0036] S12 Based on the processed data, use the causal association analysis method to construct the attack graph.
[0037] The S2 image neural network module extracts the attack graph to train the graph neural network, obtaining a graph neural network classification model;
[0038] Specific method:
[0039] S21 The image neural network module abstracts the attack graph into an adjacency matrix;
[0040] S22 Based on the adjacency matrix, an initial network structure of the graph neural network is built;
[0041] S23 Introduce the SSA parameters into the initial network structure to train the initial network structure, obtaining the graph neural network classification model.
[0042] Specific method:
[0043] S231 Set the model Loss as the fitness function of SSA, and specify the parameter optimization range to obtain the SSA parameters;
[0044] S232 Introduce the SSA parameters into the initial network structure and train the initial network structure to obtain the graph neural network classification model.
[0045] Specifically, use the SSA iterative optimization to select the sparrow individual with the lowest fitness as the best parameter combination of the model.
[0046] S3 Identify the test data through the graph neural network classification model to obtain the attack scenario.
[0047] The purpose of the present invention is to improve the efficiency of alarm correlation analysis and at the same time improve the inference ability of the model. In order to illustrate the effectiveness of the correlation analysis method in this article, first use the causal correlation analysis method to construct an attack graph, then build SSA-GCN and SSA-GAT models to learn the attack graph, and then identify the attack scenario to which the alarm belongs, and analyze the experimental results from the following aspects.
[0048] I. Compare the classification performances of different classifiers
[0049] The present invention implements random forest, SVM, and Multilayer Perceptron (MLP) with the help of the Scikit-learn machine learning library, and adjusts the best parameter combination as much as possible. The classification performance is also measured using classification evaluation metrics. As shown in the performance comparison results of the method of the present invention for different classifiers, the method of the present invention exceeds 90% in all indicators, showing a significant improvement compared with other methods, especially in terms of accuracy and recall rate, indicating that the method of the present invention generates fewer missed alarm samples and will classify alarms into the correct categories as much as possible. At the same time, the precision rate generated by the method of the present invention reaches 92.4%, indicating that the model of the present invention avoids misclassifying alarm categories, so the generated false alarm samples are also very few, as Figure 1 shown.
[0050] II. Comparing the classification performance of different graph neural networks
[0051] In order to further verify that GNN can effectively utilize the information in the alarm-related attack graph, on the one hand, the Alert-GCN classification model based on attribute similarity is experimentally compared; on the other hand, on the premise that both use the causal association analysis attack graph, the present invention also compares the performance of GCN and GAT. First, the Loss convergence curves trained by Alert-GCN and the method of the present invention are given, as Figure 2 shown. As shown in the figure, the Loss of both models converges rapidly in the first 80 epochs. After that, it slowly iterates until about 250 epochs, and the model tends to be stable. Finally, it can be seen that the Loss trained by the method of the present invention is significantly lower than that of the Alert-GCN model, and finally drops below 0.2, indicating that the method of the present invention has a better fitting degree for the attack scenario than Alert-GCN. Although the Loss value of Alert-GCN is lower in the early stage, the decline rates of the two Loss values are almost the same, and Alert-GCN starts to converge smoothly at about 100 epochs, while the Loss of the method of the present invention is still decreasing, indicating that the alarm-related attack graph plays a role at this time, and GCN can effectively utilize the information in the attack graph for inference learning.
[0052] The resources consumed for training the above two models are also different. The table shows the time, the number of edges in the relationship graph, and the accuracy rate consumed by the method of the present invention and Alert-GCN for training the model. It can be seen that for the same 500 epochs of training, the time required for GCN trained based on the attack graph is less than that of Alert-GCN, and the time efficiency is improved by about 7%. The accuracy rate of the model is improved by about 6%. Moreover, the scale of the attack graph obtained through association analysis is also smaller than that of the relationship graph of Alert-GCN. Therefore, the method of the present invention obtains better performance using less information, indicating that GCN can effectively utilize the information in the alarm-related attack graph.
[0053]
[0054] The above experiments illustrate that the attack graph obtained by association analysis is useful for attack scenario recognition. The present invention also verifies the learning ability of SSA-GAT for the alarm attack graph through experiments. Before the experiments, SSA is used to find parameters for different graph neural networks respectively, and the optimal parameter combinations of each are selected as much as possible for comparison. Before the experiments, SSA is used to find parameters for different graph neural networks respectively, and the optimal parameter combinations of each are selected as much as possible for comparison, Figure 3 The performance comparison of alarm association analysis methods based on different graph neural networks is given.
[0055] The graph structures learned by SSA-GCN and SSA-GAT both adopt the same causal association analysis attack graph. From Figure 3 it can be seen that the classification performances of SSA-GCN and SSA-GAT using the attack graph are basically better than Alert-GCN. The performance of SSA-GCN is better, with the accuracy and recall rate reaching 0.9386, and the f1 coefficient can also reach 0.92, indicating that it has good performance in all aspects of indicators. During the experiment, although SSA-GAT also has good classification ability and the recall rate can exceed 0.92, the training resources consumed by SSA-GAT are higher than those of SSA-GCN. The reason is that SSA-GAT performs an additional fully connected operation when calculating the attention coefficient and matrix splicing, resulting in the need for a large amount of hardware resources to support the training of a large number of samples.
[0056] III. Comparing the optimization performances of different parameter optimization algorithms
[0057] To verify the effectiveness of the SSA parameter optimization method adopted by the present invention, a comparison is made with the Particle Swarm Optimization (PSO). Both search for the parameter combinations of GCN. It can be seen from the table that on the premise that the fitness is almost the same, the SSA parameter optimization method adopted by the present invention has a faster convergence speed than the PSO parameter optimization method.
[0058]
[0059]
[0060] The results show that there is little difference in the convergence accuracy between SSA and PSO, and the fitness of the parameter models found by both is around 0.14. However, the convergence speed of SSA is faster than that of PSO. It only takes 23 cycles to find the optimal parameter model. The reason is that sparrow individuals with different identities can cooperate in the search. The discoverer and alarm mechanisms in SSA improve the global exploration ability, while the followers can quickly converge near the optimal value. The parameter search process of the neural network is carried out in the discrete solution set space, and sometimes PSO cannot accurately search for the global optimal solution, so the convergence speed of PSO is slower.
[0061] According to the above design framework diagram and flowchart, the alarm aggregation efficiency of the proposed method was verified on the DARPA2000 dataset. By replaying the tcpdump traffic data of the LLDOS1.0 scenario in this dataset and extracting 30,459 alarm samples based on snort, each sample has 13-dimensional features, including 10-dimensional nominal attributes and 3-dimensional numerical attributes, and their distributions are shown in the following table:
[0062]
[0063] Based on the above dataset, the specific steps of alarm correlation analysis are as follows:
[0064] Data preprocessing stage. Clean the alarm data such as standardization, and use the causal correlation analysis method to generate the alarm attack graph and attack scenarios in this stage. Otherwise, the GNN cannot be trained.
[0065] The causal correlation analysis method assumes that the later-occurring alarms are due to the successful intrusion of the previous alarms. Based on this assumption, super-alarm instances are constructed for the known alarm types, and then the following formula is used for rule matching to determine whether there is a logical relationship between the instances.
[0066] I = C(T) ∩ P(T')
[0067] where T and T' are two super-alarm instances, C(T) represents the set of result predicates that T may generate, P(T') represents the logical combination of the premise predicates of T', and it is required that the end time of C(T) is before the start time of P(T'), that is, C(T).end_time ≤ P(T').begin_time. If S is the pre-constructed attack scenario, then it means that T and T' are causally related.
[0068] The implementation steps of the causal correlation analysis method are as follows:
[0069] (1) Divide the alarm T into the appropriate attack scenario and search for the result set C(T) of T in S;
[0070] (2) In the subsequent attack phase of T, find an alarm T′ such that the causal association condition of formula I = C(T) ∩ P(T') is satisfied between C(T) and P(T′), then T is associated with T′;
[0071] (3) Set T′ as the current alarm, continue to associate backward, and repeat steps (2) and (3) until there is no subsequent alarm, then the causal association regarding alarm T ends.
[0072] 2. Training phase. Abstract the attack graph into an adjacency matrix, build the initial network structure of the GNN, set the model Loss as the fitness function of SSA, specify the parameter optimization range, use SSA for iterative optimization, and select the sparrow individual with the lowest fitness as the best parameter combination of the model.
[0073] Although the GNN can learn richer information, the hyperparameters in the network model are complex and there is a problem of parameter sensitivity. It is obviously not intelligent enough to adjust the parameters completely based on human experience. In this chapter, SSA is used for parameter optimization to train a more accurate GNN model. In this paper, two network structures, SSA - GCN and SSA - GAT, are built to train the classification model. Taking GCN as an example, the built SSA - GCN structure is as Figure 4 shown:
[0074] (1) Problem encoding. Initially build a GCN structure with 3 hidden layers. The ReLU activation function is used in all hidden layers. Use SSA to first find the number of neurons in each hidden layer, the learning rate, and the number of iteration cycles. Therefore, the dimension of the sparrow individual of SSA is set to 5 - dimensional. Assume that the number of individuals in the sparrow population X is n, then there is:
[0075]
[0076] (2) Determination of the fitness function. The purpose of SSA is to find the parameter model with the best fitting degree. Therefore, the loss function of GCN can be used as the fitness function of SSA, and the minimum loss value is searched for in each round of iteration. The GCN structure built in this chapter uses cross - entropy as the loss function. Therefore, the fitness function of SSA is:
[0077]
[0078] Among them, J represents the number of classifications, p i represents the predicted probability of the i - th class, and q i represents the true probability of this class.
[0079] 3. Testing phase. Input the alarm data to be tested into the trained GNN classification model to obtain the attack scenario to which the alarm belongs, and complete the alarm association.
[0080] Experimental results:
[0081] The present invention also verified the learning ability of SSA-GAT for the alarm attack graph through experiments. Before the experiments, SSA was used to find parameters for different graph neural networks respectively, and the optimal parameter combinations were selected as much as possible for comparison. Figure 5 The performance comparison of the alarm correlation analysis methods based on different graph neural networks is given.
[0082] The graph structures learned by SSA-GCN and SSA-GAT both adopt the same causal correlation analysis attack graph. Figure 5 It can be seen that the classification performance of SSA-GCN and SSA-GAT using the attack graph is basically better than that of Alert-GCN. The performance of SSA-GCN is better, with the accuracy and recall rate reaching 0.9386, and the f1 coefficient can also reach 0.92, indicating that it has good performance in all aspects. From the above, it can be seen that the accuracy of SSA-GCN is significantly higher than that of SSA-GAT and Alert-GCN. During the experiment, although SSA-GAT also has good classification ability and the recall rate can exceed 0.92, the training resources consumed by SSA-GAT are higher than those of SSA-GCN. The reason is that SSA-GAT performs an additional fully connected operation when calculating the attention coefficient and matrix splicing, resulting in the need for a large amount of hardware resources to support the training of a large number of samples.
[0083] The above-disclosed is only a preferred embodiment of the present invention, and of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand the entire or part of the process of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. An alarm correlation analysis method based on a graph neural network, characterized in that It includes the following steps: Preprocess the alarm data through the causal association module to obtain an attack graph; The image neural network module extracts the training graph neural network from the attack graph to obtain a graph neural network classification model; Specific method: The image neural network module abstracts the attack graph into an adjacency matrix; Based on the adjacency matrix, build an initial network structure of the graph neural network; Introduce the SSA parameter into the initial network structure to train the initial network structure to obtain the graph neural network classification model; Specific method: Set the model Loss as the fitness function of SSA and specify the parameter optimization range to obtain the SSA parameter; Introduce the SSA parameter into the initial network structure and train the initial network structure to obtain the graph neural network classification model; Identify the test data through the graph neural network classification model to obtain an attack scenario.
2. The alarm correlation analysis method based on the graph neural network according to claim 1, characterized in that The specific method of preprocessing the alarm data through the causal association module to obtain an attack graph: The causal association module performs standardized cleaning on the alarm data to obtain processed data; Based on the processed data, use the causal association analysis method to construct the attack graph.
3. The alarm correlation analysis method based on the graph neural network according to claim 2, characterized in that The visualization processing software is the drawing software graphviz.