Security event association aggregation and maliciousness analysis method
The encoder-decoder architecture and improved DBSCAN method are used to perform security incident association aggregation and malicious analysis, which solves the problem of insufficient utilization of event context information in the prior art, realizes automatic processing of security incidents and malicious evaluation, and improves the efficiency and accuracy of security analysis.
Patent Information
- Application Number
- CN202510362271.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-04
AI Technical Summary
The original security data generated by existing security monitoring policies is redundant and trivial, making it difficult for security analysts to quickly identify useful information. The existing methods are insufficiently utilized by event context information, have poor flexibility, and are difficult to adapt to changeable attack scenarios.
The encoder-decoder architecture is used to combine the attention mechanism for event association, the improved DBSCAN method is used for clustering, and malicious analysis is performed through the LDA theme model to achieve automated processing and malicious evaluation of event sequences.
It effectively reduces the workload of manual analysis, improves the ability to identify event associations, reduces the computational complexity, and accurately recognizes the maliciousness of event sequences, reducing the false positive rate.
Smart Images

Figure CN120263452A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to a method for security event correlation aggregation and maliciousness analysis. Background Art
[0002] With the continuous progress of network technology, the types of attacks faced by enterprises are becoming more and more diverse at present, such as phishing attacks, worm attacks, denial-of-service attacks, etc., and various advanced multi-step attacks are also emerging more and more frequently. In order to better prevent and respond to various attacks, security monitoring strategies are widely used to maintain the security of IT infrastructure. By real-time monitoring data such as the running status of the system and network traffic, and further analyzing the relevant data, it is possible to more effectively protect the enterprise's information system and data and improve the response ability to security threats.
[0003] Although the existing security monitoring strategies have greatly improved the threat recognition ability, in the current complex IT system environment, the original security data generated by various security monitoring strategies are mostly trivial and redundant. Any small problem may generate a large number of security alert events in a short time, making it difficult for security analysts to quickly identify useful information from a large number of alerts. This high workload leads to a situation called "alert fatigue", that is, security analysts cannot respond in a timely and accurate manner due to the large number of security alerts received every day, and ignore the real threats hidden in the massive events. Therefore, correlating and aggregating security events to provide rich and detailed context information for events can reveal the internal connections between multiple seemingly unrelated events, which helps security analysts intuitively analyze the attack trend. Conducting maliciousness analysis on the basis of event correlation aggregation can provide a more comprehensive and in-depth security perspective, which is conducive to providing a direct basis for the decision-making of security analysts. Studying the method of security event correlation aggregation and maliciousness analysis is of great significance for improving the security protection level of enterprises or organizations, optimizing the allocation of security resources, formulating effective security policies and risk management measures.
[0004] The paper "A new alert correlation model based on similarity approach[C] / / 2019 1st International Conference on Cybernetics and Intelligent System(ICORIS).IEEE,2019,1:133-137." combines and correlates events based on similarity, and then evaluates the maliciousness of events according to the experience of historical attack events. The article describes a two-step method for low-level and high-level alerts. For low-level alerts, they determine the equality of the attributes of network flows and obtain a similarity value matrix for each alert type. For high-level alerts, they measure the similarity of the obtained alert matrix. This method can alleviate alert fatigue to a certain extent and reduce the manual burden. However, such methods often focus on the similarity of the surface features of events rather than the context of events. Insufficient utilization of context information may lead to misunderstandings of the true relationships between events. And the setting of the similarity function makes the method less flexible, restricting the automatic processing level of the method.
[0005] The paper "Alerts correlation and causal analysis for APT based cyberattack detection[J].IEEE Access,2020,8:162642-162656." uses a step-based method to reconstruct attack steps by analyzing the logical relationships and time series between events. The article proposes a real-time attack detection method based on APT. They conduct a causal analysis of the meta-alerts generated by security sensors and other system loggers. By calculating the sensitivity of each host to possible APT attacks through dynamic programming, this dynamic programming processes the alerts from each host separately and conducts a long-term analysis of the attack process. Although it performs better in understanding the relationships between events, methods that use preset rules or patterns to identify relevant events and their maliciousness often require a large amount of expert knowledge and prior information to define the logical steps and relationships between events. The implementation of some methods, such as those using attack graphs and attack trees for event correlation aggregation, is often more complex and difficult to adapt to changing attack scenarios.
[0006] In summary, the similarity-based method is simple and easy to implement. However, this method makes insufficient use of the context information of events and has very limited ability to detect the correlation between events. At the same time, the function for calculating correlation is too specific, resulting in poor flexibility of the similarity-based method. The step-based method has easily interpretable security event correlation results. However, some methods rely too much on preset rules and patterns, and the implementation is more complex, making it difficult to improve the level of automated processing. The hybrid method allows combining two or more security event correlation aggregation methods, which not only provides the ability to combine different correlation methods but also can judge the maliciousness of events by analyzing the relationships between events. Therefore, the present invention uses a hybrid method for event correlation aggregation. First, a step-based correlation strategy is used to identify the complex relationships between event contexts and form associated event sequences. Then, a similarity-based clustering strategy is applied to merge similar event sequences to reduce the workload of subsequent manual analysis and relieve alert fatigue. Furthermore, based on event correlation aggregation, maliciousness analysis is further performed on the event sequences to explore the nature of the associated events and distinguish malicious activities from benign behaviors. Summary of the Invention
[0007] The object of the present invention is to design and implement a security event correlation aggregation and maliciousness analysis method, which can automatically analyze and discover the correlation relationships between a large number of events, reduce the workload of manual analysis, and at the same time can perform maliciousness analysis on the associated event sequences to identify the nature of the event sequences.
[0008] (1) The event correlation method based on the encoder-decoder architecture combined with the attention mechanism proposed in this application transforms the problem of discovering the correlation relationships of security events into a sequence-to-sequence problem. Through training, the model can generate correct predictions, and then the correlation between events is reflected by the attention vectors that can obtain correct predictions. (2) An event sequence aggregation method based on improved DBSCAN proposed in this application clusters event sequences with similar context features, thereby reducing the data complexity and computational cost of subsequent analysis. (3) A maliciousness analysis method for event sequences is proposed based on the LDA topic model in this application, which can better understand the potential intentions behind the associated events.
[0009] The technical solution of the present invention is as follows: A security event correlation aggregation and maliciousness analysis method includes the following steps:
[0010] Preprocess the security event log set to obtain the security event context sequence and its embedding vector; the embedding vector of the security event context sequence is obtained as the initial vector C by the encoder; the decoder calculates the attention vector and the predicted event probability distribution according to the embedding vector decoded in the previous step and the initial vector C.
[0011] After iterative training meets the termination condition, calculate the total attention distribution of each event according to the obtained attention vector and the security event context sequence; use the improved DBSCAN method to cluster the dataset Z composed of the total attention distributions of all events; extract samples from each cluster and perform event maliciousness analysis according to the LDA topic model.
[0012] The process of obtaining the security event context sequence and its embedding vector is as follows:
[0013] Step 1.1: Use a log parsing tool to convert unstructured original log information into a structured log template.
[0014] Step 1.2: Obtain the structured event template of each log record by splitting compound words and removing special symbols, and assign a unique event category identifier id to each structured event template.
[0015] Step 1.3: Group the security events according to device information and sort the security events according to the timestamp.
[0016] Step 1.4: Determine the optimal values of the maximum length L and the maximum time difference t of the security event context sequence through experiments, and obtain each security event and its corresponding security event context sequence: (IDs, id y ), where id y represents the category of the security event, and IDs = {id1, id2,..., id L} represents the category of each event in the security event context sequence corresponding to this security event. When the time difference between the event in the security event context sequence and this event exceeds t, set the corresponding id to -1.
[0017] Step 1.5: Use the BERT model to convert the security event context sequence IDs into a vector of n_dim dimensions. batchsize represents the amount of data processed each time, and L represents the maximum length of the security event context sequence:
[0018] Word embedding layer vector: T ∈ R (batchsize,L,n_dim)
[0019] Paragraph embedding layer vector: S ∈ R (batchsize,L,n_dim)
[0020] Position embedding layer vector: P ∈ R (batchsize,L,n_dim)
[0021] Embedding vector of the security event context sequence: X = T + S + P.
[0022] Among them, a security event is just a general term for events in the log and does not represent the security of the event itself. For example, there are 3 records in the log, and each record corresponds to a security event, as follows:
[0023] Device A of e1 requests to query database D
[0024] Device B of e2 requests to query database C
[0025] Device A of e3 requests to access the file directory of device B
[0026] Step 1.2 is to extract the template of each security event according to the event content:
[0027] The content structures of e1 and e2 are similar and belong to 1 structured event template, and the event category identifier id is set to 1 for both;
[0028] e3 belongs to another template, and the time category identifier id is 2.
[0029] The encoder is of CNN-BiGRU structure, including a CNN network, a BiGRU network and a fully connected layer; the embedding vectors of the security event context sequence respectively obtain spatial feature extraction data through the CNN network and time series feature data through the BiGRU network; the spatial feature extraction data and the time series feature data are merged and then reduced in dimension through the fully connected layer, and are regularized through dropout and encoded into an initial vector C containing context information.
[0030] The decoder is of GRU-Attention structure, including an Attention module and an event weighted calculation module. The Attention module consists of an embedding layer, a GRU, a linear layer and a softmax layer; the event weighted calculation module consists of a matrix multiplication layer, a linear layer and a softmax layer; the Attention module inputs the embedding vector of the predicted event probability distribution e i-1 from the previous step of the decoder and the initial vector C into the GRU to calculate the attention score and update the initial vector C. The attention score is normalized through the linear layer and the softmax layer of the Attention module to be converted into an attention vector; the embedding vector of the security event context sequence input to the encoder and the attention vector α = {α1, α1, …, α L} obtained by the Attention module are input to the matrix multiplication layer of the event weighted calculation module to perform element-wise multiplication, weight the embedding vectors in the security event context sequence, and convert the weighted security event vector through the linear layer and the softmax layer of the event weighted calculation module into a predicted event probability distribution.
[0031] The termination condition of the iterative training is to reach the maximum number of iterations or the total loss is lower than the target value;
[0032] Total loss function:
[0033] Among them, y represents the true probability distribution of security events, represents the event probability distribution predicted by the encoder-decoder according to the security event context sequence, N represents the number of event types, and P(y j ) represents the probability that the true event is the j-th type of event, represents the probability that the predicted event is the j-th type of event;
[0034]
[0035] Among them, Softmax(·) represents the softmax function, R(·) represents the ReLU activation function, W(·)+b represents the linear transformation layer, W represents the weight matrix, b represents the bias term, and α i is the attention weight corresponding to the i-th security event in the security event context sequence, and x i represents the embedding vector of the i-th security event in the security event context sequence.
[0036] The total attention distribution of each event is calculated as follows:
[0037] The category of each event in the security event context sequence: IDs = {id1, id2,..., id L}
[0038] The total attention distribution of events: V = {v1, v2,..., v N}
[0039]
[0040] Among them, 1[id j = i] is an indicator function that takes the value of 1 when the category of the j-th event in the security event context sequence is equal to i, and 0 otherwise; id j represents the category of the j-th security event in the security event context sequence.
[0041] The improved DBSCAN method is as follows;
[0042] Use the K-means algorithm to partition the dataset Z into k non-overlapping partitions;
[0043] In each partition generated by the K-means algorithm, first randomly select an event to join the seed event set Y, and then use the farthest sampling method based on the Manhattan distance to select num i· r seed events are added to Y, and the initially randomly selected events are removed from Y after sampling; where num i represents the number of events in the i-th partition, and r represents the sampling ratio of the seed events; the DBSCAN clustering algorithm is used on the seed event set Y; according to the results of DBSCAN clustering on the seed event set Y, each event in the data set Z is checked; if the event is in the seed event set Y and is assigned to a certain cluster, the same cluster is also assigned in the original data set Z; if it is in the seed event set Y but is identified as noise, it is also marked as noise in the original data set Z; if it is not in the seed event set but is density-reachable from a certain core point, it is classified into the corresponding cluster; if it is not in the seed event sequence and is not density-reachable from all core points, it is judged whether it is density-reachable from a certain core point. If it is density-reachable, it is classified into the corresponding cluster, otherwise it is marked as noise, so as to obtain the final clustering result Z ′ 。
[0044] Furthermore, in order to balance the performance and quality of clustering, the sampling ratio of the seed event sequence is set to 40% according to experience.
[0045] The maliciousness analysis is as follows:
[0046] Step 4.1: First, represent the security events in the security event context sequence in the form of a triple: e=(id, t, event), where id represents the unique identifier of the event, t represents the timestamp when the event occurs, and event describes the name of the event;
[0047] Step 4.2: According to the final clustering result Z ′ , n security event context sequences are extracted from each cluster as samples by the farthest sampling method based on distance; if the number of event sequences in the cluster is less than n, all the security event context sequences in the cluster are used as samples;
[0048] Step 4.3: Apply the LDA topic model to the set of sample security event context sequences in each cluster, and approximately estimate the model parameters through the Gibbs sampling process to obtain the converged security event context sequence - topic distribution θ and topic - event distribution where the set of sample security event context sequences is regarded as a document set, each single event sequence in the set of sample security event context sequences is equivalent to a document, and each single event in the security event context sequence corresponds to a word in the document;
[0049] Step 4.4: According to the security event context sequence - topic distribution θ, use the topic distribution probability of each security event context sequence as the score of the topic;
[0050] Step 4.5: According to the five-stage attack model, the attack is divided into five stages: initial intrusion, establishing a foothold, lateral movement, penetration / hindrance, and post-penetration / post-hindrance. A score is assigned to each attack stage based on the importance of each stage and the severity of the consequences. When an event can be classified into multiple stages, the stage with the highest score is selected. Combining expert knowledge and experience, according to the correlation between different task categories and keyword events in the dataset and different attack stages, the events are assigned to one of the stages to obtain a maliciousness rating table for the dataset.
[0051] Step 4.6: According to the topic-event distribution Select the top top_n events with the highest probability from each topic-event distribution; if the event sequence contains the above events, match these events with the maliciousness rating table and add the scores of the corresponding levels. Multiple events of the same level are only counted once. Calculate the score s of the event sequence on each topic by accumulating the scores. k The sum of the scores of each topic is used to evaluate the correlation between the topic and malicious attacks; then, normalize the scores of each topic to obtain the weight ω of the event sequence for each topic. k The calculation formula is as follows:
[0052]
[0053] Step 4.7: For each sample event sequence, calculate its maliciousness score p mal If it exceeds the preset threshold p0, then determine that the event sequence is a malicious attack; the calculation formula is:
[0054]
[0055] Among them, θ k represents the probability that the event sequence belongs to the k-th topic in the event sequence-topic distribution θ in Step 4.4.
[0056] Step 4.8: Finally, based on the malicious attack determination results of the sample event sequences, estimate the maliciousness of the entire clustering; if all sample event sequences are determined to be malicious attack sequences, then consider that the entire clustering to which it belongs belongs to malicious attacks; if no sample event sequence is determined to be a malicious attack sequence, then consider that the entire clustering to which it belongs does not belong to malicious attacks; if there are some sample event sequences determined to be malicious attack sequences, it means that the clustering contains both malicious attack sequences and normal behavior sequences. Security analysts need to further analyze the reasons, mark the clustering as ambiguous, and re-execute Steps 4.3 to 4.7 for all sample event sequences in the clustering.
[0057] Furthermore, the log parsing tool uses Drain.
[0058] Advantages of the present invention: In view of the problems existing in the existing correlation methods to varying degrees, such as insufficient attention to event context, low level of event automation processing, and failure to further analyze the maliciousness of events, the present invention proposes a method for security event correlation aggregation and maliciousness analysis. The key points mainly include the following three points:
[0059] (1) By adopting an encoder-decoder model combined with an attention mechanism, using the BERT model for word embedding, and using parallel CNN and BiGRU as encoders, and GRU combined with an attention mechanism as a decoder, the dynamic evaluation and automated processing of event correlation relationships are realized.
[0060] (2) A method for aggregating event sequences based on improved DBSCAN is proposed. By using the Manhattan distance as the distance metric, the workload that needs to be further analyzed can be effectively reduced.
[0061] (3) A method for maliciousness analysis based on the LDA topic model is proposed. Regarding the set of sample event sequences as a document set, a single event sequence as a document, and a single event as a word, the "event sequence - topic" distribution and "topic - event" distribution are obtained using the LDA topic model to score the event sequences, thereby analyzing the maliciousness of the event sequences. Description of the Drawings
[0062] Figure 1 It is a flowchart of a method for security event correlation aggregation and maliciousness analysis;
[0063] Figure 2 It is a schematic diagram of the processes of the encoder, decoder, and LDA topic model;
[0064] Figure 3 It is a schematic diagram of the maliciousness analysis of event sequences based on the LDA topic model;
[0065] Figure 4 It is the prediction accuracy of events with the maximum number of events in different contexts;
[0066] Figure 5 It is the prediction accuracy of events with the maximum time gap in different contexts;
[0067] Figure 6 It is the prediction accuracy of events with different models;
[0068] Figure 7 It is the comparison of the running times of improved DBSCAN and standard DBSCAN; Detailed Embodiment
[0069] The following further describes the specific implementation manners of the present invention in conjunction with the accompanying drawings and implementation examples.
[0070] In this implementation example, the method proposed by the present invention is comprehensively evaluated using the HDFS log dataset and the simulation log dataset.
[0071] The HDFS log dataset is a publicly available dataset generated by running the Hadoop-based MapReduce framework on 203 Amazon EC2 nodes and manually labeled by Hadoop domain experts. Among the approximately 11.2 million log entries collected, there are approximately 2.9% abnormal events. Since there are no relevant labels for multi-step attacks in the HDFS dataset, the HDFS dataset is mainly used for event prediction and verification in event reduction.
[0072] The simulation log dataset is a log dataset reconstructed based on the log dataset generated in the paper "Alsaheel A, Nan Y, Ma S, et al. ATLAS: A sequence-based learning approach for attack investigation [C] / / 30th USENIX security symposium (USENIX security 21). 2021: 3005 3022."
[0073] First, use the attack graph file obtained by the ATLAS model to restore the main steps of each single-host attack, and then extract the core log records related to multi-step attacks from the original dataset as the atomic attack events of multi-step attacks and mark them. Then, set two threads to replay the atomic events of multi-step attacks and benign activity events respectively, so as to obtain a simulation log dataset of 4 single-host attacks. The information of the 4 single-host attacks is shown in Table 1.
[0074] Table 1 Simulation log dataset of 4 single-host attacks
[0075] Attack ID Attack activity Exploiting CVE vulnerability Size (MB) Number of events S-1 Strategic web 2015-5122 381 95.0K S-2 Malvertising dominate 2015-3105 990 397.9K S-3 Spam campaign 2017-11882 521 128.3K S-4 Pony campaign 2017-0199 448 125.6K
[0076] A method for security event correlation aggregation and maliciousness analysis, the specific steps are as follows:
[0077] Step 1: Perform preprocessing operations on the security event log set to obtain the security event context sequence and its embedding vector. The specific steps are as follows:
[0078] Use the log parsing tool Drain to parse the security event log and convert the unstructured original log information into a more easily processed structured log template.
[0079] Perform a series of optimization operations on the security event data to further optimize the data structure. This process includes, but is not limited to, steps such as splitting compound words and removing special symbols, to obtain the structured event template for each log record.
[0080] Group the security events according to device information and sort the security events based on the timestamp to ensure the organizational structure and time continuity of the data.
[0081] Select the data that ranks in the top 5% in terms of time sorting from the two datasets, determine the sliding window that limits the context event range and the maximum time gap between the first event and the last event within the window. These data do not participate in the subsequent training process. Use 80% of the above data as the training set and 20% as the test set, and test the prediction accuracy of the security event context sequence for the maximum lengths L with the maximum number being 2, 5, 8, 10, 20, 25, 30 and the maximum time gaps t being 1 minute, 10 minutes, half an hour, 1 hour, 1 day, 1 week respectively. The experimental results are respectively as Figure 4 and Figure 5 shown. For the HDFS dataset, the maximum number of the security event context sequence is 10, and the maximum time gap is determined to be 1 day; for the simulation log dataset, the maximum number of the security event context sequence is 20, and the maximum time gap is 10 minutes.
[0082] Use the BERT model to convert the security event into a 768-dimensional vector:
[0083] Word embedding layer vector: T ∈ R (batchsize,L,768)
[0084] Paragraph embedding layer vector: S ∈ R (batchsize,L,768)
[0085] Position embedding layer vector: P ∈ R (batchsize,L,768)
[0086] Comprehensive embedding vector of the security event: X = T + S + P
[0087] Step 2: Use the event association method based on the encoder-decoder architecture to associate the embedding vectors of the security event context sequence
[0088] Input the embedding vectors of the security event context sequence into the encoder of CNN-BiGRU, and merge the spatial feature extraction data obtained by CNN and the time series data obtained by BiGRU. Then, reduce the merged high-dimensional features to 128 dimensions through a fully connected layer and use dropout for regularization to prevent overfitting to obtain the input of the decoder.
[0089] Decode using a GRU-Attention based decoder, which obtains the embedding vectors of the security event context sequence, performs element-wise multiplication with the attention vectors obtained by the attention layer to weight these events, and predicts the probability distribution of the events through a neural network with a 128-dimensional linear layer equipped with ReLU activation function and softmax function.
[0090] Evaluate the model performance by comparing the "predicted" probability distribution with the actual events. The present invention uses KL divergence as the loss function for backpropagation and the Adam optimization algorithm to adjust the attention weights and other parameters in the model to find the optimal solution during training:
[0091] Computation process of the decoder:
[0092] Loss function:
[0093] Initial learning rate: θ = 0.01
[0094] Decay rate of the first moment estimate: β1 = 0.9
[0095] Decay rate of the second moment estimate: γ2 = 0.98
[0096] Minimum value to prevent division by zero: ε = 1e -9
[0097] Step 3: Use the improved DBSCAN clustering algorithm to aggregate the attention-weighted event sequences, simplify a large number of event sequences into several clusters that are easier to manage and analyze, significantly reduce the data complexity, and at the same time retain the key information and structural features of the event sequences.
[0098] Use the K-means algorithm to partition the original event sequence dataset into k non-overlapping partitions.
[0099] In each partition generated by the K-means algorithm, further perform sampling of the seed event sequences to form a new and more representative dataset. To balance the performance and quality of clustering, the sampling ratio of the seed event sequences is set to 40% according to experience, providing high-quality input data for the subsequent application of the DBSCAN clustering algorithm.
[0100] Use the DBSCAN clustering algorithm on the newly formed event sequence dataset, where ε is set to 0.2 and MinPts is set to 5.
[0101] Map the results of DBSCAN clustering to the original event sequence dataset. By examining each sequence in the original event sequence set, if it is in the seed event sequence set and is assigned to a certain cluster, then the same cluster is also assigned in the original event sequence set; if it is in the seed event sequence set but is identified as noise, then it is also marked as noise in the original event sequence set; if it is not in the seed event sequence set but is density-reachable from a certain core point, it is classified into the corresponding cluster; if it is not in the seed event sequence and is far from all core points, then determine whether it is density-reachable from a certain core point. If it is density-reachable, it is classified into the corresponding cluster, otherwise it is marked as noise, thus obtaining the final clustering result.
[0102] Set different values for k respectively, calculate the average silhouette coefficient, and select the k value with a relatively high average silhouette coefficient as the number of pre-partitions.
[0103] Step 4: Use the LDA topic model to analyze the maliciousness of event sequences.
[0104] First, format the event sequences, and represent the original events in the event sequences in the form of triples: e=(id, t, event), where id represents the unique identifier of the event, t represents the timestamp when the event occurs, and event describes the name of the event.
[0105] In each cluster, extract n (set n = 5) event sequences as samples by the farthest sampling method based on distance. If the number of event sequences in the cluster is less than n, then all event sequences in the cluster are used as samples.
[0106] Apply the LDA topic model to the sample event sequence set in each cluster, and approximately estimate the model parameters through the Gibbs sampling process to obtain the converged security event context sequence - topic distribution θ and topic - event distribution Among them, the sample event sequence set is regarded as a document set, each single event sequence in the set is equivalent to a document, and a single event in the event sequence corresponds to a word in the document. Since the present invention aims to analyze the maliciousness of event sequences, the number of topics of the LDA topic model is set to two, that is, K = 2.
[0107] According to the event sequence - topic distribution θ k , take the topic distribution probability of each event sequence as the score of the topic.
[0108] According to the five-stage attack model commonly used in most literature, a score is assigned to each attack stage based on the importance of each stage and the severity of the possible consequences. When an event can be classified into multiple stages, the stage with the highest score is selected. Then, combining expert knowledge and experience, according to the different task categories in the dataset and the relevance of the events of keywords to different attack stages, the events are assigned to one of the stages to obtain a maliciousness rating table for the dataset.
[0109] According to the theme-event distribution Select the top 10 events with the highest probabilities from each theme-event distribution. If the event sequence contains the above events, these events are matched with the maliciousness rating table and the scores of the corresponding levels are added. Multiple events of the same level are only scored once. The score s of the event sequence on each theme is calculated by accumulating the scores. k (k = 1, 2), the sum of the scores of each theme is used to evaluate the degree of correlation between the theme and malicious attacks. Then, the scores of each theme are normalized to obtain the weight ω of the event sequence for each theme. k , and the calculation formula is as follows:
[0110]
[0111] For each sample event sequence, calculate its maliciousness score p mal , if it exceeds the preset threshold p0 = 0.5, then the event sequence is determined to be a malicious attack. The calculation formula is:
[0112]
[0113] Finally, according to the malicious attack determination results of the sample event sequences, the maliciousness of the entire clustering is estimated. If all sample event sequences are determined to be malicious attack sequences, then the entire clustering to which they belong is considered to be malicious attacks. If no sample event sequence is determined to be a malicious attack sequence, then the entire clustering to which they belong is considered not to be malicious attacks. If there are some sample event sequences determined to be malicious attack sequences, it means that the clustering may contain both malicious attack sequences and normal behavior sequences, and security analysts need to further analyze the reasons, mark the clustering as ambiguous, and re-execute all the sample event sequences in the clustering.
[0114] To verify the effectiveness of the method proposed in the present invention, a set of experiments will be carried out from four aspects: event prediction and association, clustering performance, event reduction, and maliciousness analysis effect, and compared with different methods to verify the effectiveness of the method of the present invention.
[0115] We set 80% of the data in the dataset for training and 20% of the data for testing.
[0116] In terms of event correlation and prediction, in order to verify the effectiveness of this method in event prediction, it is compared with the following four encoder-decoder models combined with the attention mechanism as the baseline method: CNN encoder-GRU decoder, BiGRU encoder-GRU decoder, encoder-GRU decoder with serial CNN and BiGRU, and the model of this method using one-hot encoding. By Figure 6 It can be seen that this method has the highest event prediction accuracy, indicating that the proposed security event correlation method of the present invention has the ability to identify complex relationships between events. In order to verify the correlation effect of this method, this method is compared with the above baseline models and the method of event correlation based on the causal relationship graph, and experiments are carried out on the simulation log dataset. The results are shown in Table 2. It can be seen that the accuracy, recall rate and F1 score of this method are all better than other methods.
[0117] Table 2 Comparison of correlation effects of different models
[0118] Method Precision Recall F1-score CNN-GRU 84.70% 88.57% 86.59% BiGRU-GRU 85.33% 89.71% 87.47% CNN+BiGRU-GRU 86.49% 91.43% 88.89% One-hot 85.00% 87.43% 86.20% Correlation method based on causal relationship graph 84.09% 84.57% 84.33% The present invention 88.03% 89.95% 89.01%
[0119] In terms of clustering performance, in order to verify the clustering efficiency and clustering quality of the event sequence aggregation method of the present invention, the improved DBSCAN algorithm for determining the number of pre-partitions is compared with the standard DBSCAN algorithm, and it is run 5 times on each dataset to calculate the average running time. The results are as Figure 7 shown. The results show that the improved DBSCAN has improved in terms of computational efficiency compared to the standard DBSCAN.
[0120] In terms of event reduction, in order to evaluate the performance of the proposed security event correlation aggregation method of the present invention in reducing the workload of further analysis, this method is compared with the correlation method based on the causal relationship graph, the context correlation in this method combined with the DBSCAN method, and the method of directly aggregating without context correlation. The results are shown in Table 3. It can be seen that this method has achieved the best overall event reduction rate on two datasets compared with other methods, indicating that this method can effectively reduce the workload that needs to be further analyzed.
[0121] Table 3 Comparison of event reduction effects of different methods
[0122]
[0123]
[0124] In terms of maliciousness analysis, since the HDFS dataset lacks relevant markings for multi-step attack scenarios, the maliciousness analysis of the event sequence is only verified through the simulation log dataset. To verify the recognition effect of the method on 4 types of single-host attacks in the simulation log dataset, the method with the threshold p0 = 0.5 is compared with ATLAS, and the results are shown in Table 4. It can be seen that the F1 scores of the method in detecting 4 types of single-host attacks are slightly lower than those of ATLAS, but the difference in F1 scores from ATLAS is not significant. Moreover, the recall rate of the method in detecting 4 types of single-host attacks is higher than that of ATLAS, indicating that this method can capture most of the attack sequences in the 4 types of single-host attacks and avoid false negatives as much as possible. The relatively low precision indicates that the method mispredicts benign event sequences as attack behaviors, which may be caused by overestimating the score of a single event. Before semantic recognition, this method has reduced most of the events that need to be manually analyzed through the correlation aggregation method, and although there is a certain error in the precision compared with ATLAS, the difference is small. Therefore, the impact of false positives on the response efficiency is limited, indicating that this method can identify the maliciousness behind the event sequence to a certain extent.
[0125] Table 4 Comparison of maliciousness analysis effects of different methods
[0126]
[0127]
[0128] The above preferred embodiments are only for illustrating the technical concept and features of the present invention, aiming to enable those skilled in the art to understand the content of the present invention and implement it, and shall not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for security event correlation aggregation and maliciousness analysis, characterized in that, The steps are as follows: Preprocess the security event log set to obtain the security event context sequence and its embedding vector; the embedding vector of the security event context sequence is encoded by an encoder to obtain the initial vector C; The decoder calculates the attention vector and the predicted event probability distribution according to the embedding vector decoded in the previous step and the initial vector C; After iterative training until the termination condition is met, calculate the total attention distribution of each event according to the obtained attention vector and the security event context sequence; Adopt an improved DBSCAN method to cluster the data set Z composed of the total attention distributions of all events; extract samples from each cluster and perform event maliciousness analysis according to the LDA topic model.
2. The security event correlation aggregation and maliciousness analysis method according to claim 1, wherein The process of obtaining the security event context sequence and its embedding vector is as follows: Step 1.1: Use a log parsing tool to convert unstructured original log information into a structured log template; Step 1.2: Obtain the structured event template of each log record by splitting compound words and removing special symbols, and assign a unique event category identifier id to each structured event template; Step 1.3: Group the security events according to device information and sort the security events according to the time stamp; Step 1.
4. Determine the optimal values of the maximum length L of the security event context sequence and the maximum time difference t through experiments, and obtain each security event and its corresponding security event context sequence: (IDs, id y ), where id y represents the category of the security event, and IDs = {id1, id2, …, id L} represents the category of each event in the security event context sequence corresponding to this security event. When the time difference between an event in the security event context sequence and this event exceeds t, set the corresponding id to -1; Step 1.5: Use the BERT model to convert the security event context sequence IDs obtained in Step 1.4 into vectors of n_dim dimensions, where batchsize represents the amount of data processed each time, and L represents the maximum length of the security event context sequence: Word embedding layer vector: T ∈ R (batchsize,L,n_dim) Paragraph embedding layer vector: S ∈ R (batchsize,L,n_dim) Position embedding layer vector: O ∈ R (batchsize,L,n_dim) Embedding vector of the security event context sequence: X = T + s + P.
3. The security event association aggregation and maliciousness analysis method according to claim 1, characterized in that The encoder is a CNN-BiGRU structure, including a CNN network, a BiGRU network and a fully connected layer; the embedding vector of the security event context sequence obtains spatial feature extraction data through the CNN network and time series feature data through the BiGRU network; the spatial feature extraction data and the time series feature data are merged and then reduced in dimension through the fully connected layer, and regularized through dropout and encoded into an initial vector C containing context information.
4. The security event association aggregation and maliciousness analysis method according to claim 3, wherein The decoder is of the GRU-Attention structure, including an Attention module and an event weighted calculation module. The Attention module consists of an embedding layer, a GRU, a linear layer, and a softmax layer; the event weighted calculation module consists of a matrix multiplication layer, a linear layer, and a softmax layer; the Attention module inputs the embedding vector and the initial vector C of the predicted event probability distribution e output by the previous-step decoder into the GRU to calculate the attention scores and update the initial vector C, and the attention scores are normalized and transformed into an attention vector through the linear layer and the softmax layer of the Attention module; the embedding vector of the security event context sequence input to the encoder and the attention vector α = {α1, α1, …, α i-1} obtained from the Attention module are input to the matrix multiplication layer of the event weighted calculation module to perform element-wise multiplication, weight the embedding vectors in the security event context sequence, and the weighted security event vector is passed through the linear layer and the softmax layer of the event weighted calculation module to be transformed into a predicted event probability distribution. L} 5. The security event correlation aggregation and maliciousness analysis method according to claim 4, characterized in that, The termination condition of the iterative training is to reach the maximum number of iterations or the total loss is lower than the target value; Total loss function: Among them, y represents the true probability distribution of security events, represents the event probability distribution predicted by the encoder-decoder based on the security event context sequence, N represents the number of event types, P(y j ) represents the probability that the true event is the j-th type of event, represents the probability that the predicted event is the j-th type of event; Among them, Softmax(·) represents the softmax function, R(·) represents the ReLU activation function, W(·) + b represents the linear transformation layer, W represents the weight matrix, b represents the bias term, and α i is the attention weight corresponding to the i-th security event in the security event context sequence, and x i represents the embedding vector of the i-th security event in the security event context sequence.
6. The security event correlation aggregation and maliciousness analysis method according to claim 5, wherein, The calculation of the total attention distribution of each event is as follows: The category of each event in the security event context sequence: IDs = {id1, id2, …, id L} Total attention distribution of events: V = {v1, v2, …, v N} where, 1[id j = i] is an indicator function that takes the value of 1 when the j-th event category of the security event context sequence is equal to i, and 0 otherwise; id j represents the category of the j-th security event in the security event context sequence.
7. The security event correlation aggregation and maliciousness analysis method according to claim 1, characterized in that The specific improved DBSCAN method is as follows; Use the K-means algorithm to partition the data set Z and divide it into k non-overlapping partitions; In each partition generated by the K-means algorithm, first randomly select an event to add to the set of seed events Y, and then use the farthest sampling method based on the Manhattan distance to select numi1·r seed events to add to Y. After sampling, remove the initially randomly selected event from Y; where num i represents the number of events in the i-th partition, and r represents the sampling ratio of the seed events; use the DBSCAN clustering algorithm on the set of seed events Y; according to the results of DBSCAN clustering on the set of seed events Y, check each event in the dataset Z; if the event is in the set of seed events Y and is assigned to a certain cluster, then assign the same cluster in the original dataset Z; if it is in the set of seed events Y but is identified as noise, then mark it as noise in the original dataset Z; if it is not in the set of seed events but is density-reachable from a certain core point, then assign it to the corresponding cluster; if it is not in the seed event sequence and is not density-reachable from all core points, determine whether it is density-reachable from a certain core point. If it is density-reachable, then assign it to the corresponding cluster, otherwise mark it as noise, so as to obtain the final clustering result Z'.
8. The security event association aggregation and maliciousness analysis method according to claim 7, wherein The maliciousness analysis is as follows: Step 4.1: First, represent the security events in the security event context sequence in the form of a triple: e = (id, t, event), where id represents the unique identifier of the event, t represents the time stamp when the event occurs, and event describes the name of the event; Step 4.2: According to the final clustering result Z′, extract n security event context sequences from each cluster as samples by the farthest sampling method based on distance; if the number of event sequences in the cluster is less than n, then use all the security event context sequences in the cluster as samples; Step 4.3: Apply the LDA topic model to the set of sample security event context sequences in each cluster, and approximately estimate the model parameters through the Gibbs sampling process to obtain the converged security event context sequence - topic distribution θ and topic - event distribution Among them, the set of sample security event context sequences is regarded as a document set. Each single event sequence in the set of sample security event context sequences is equivalent to a document, and each single event in the security event context sequence corresponds to a word in the document; Step 4.4: According to the security event context sequence - topic distribution θ, use the topic distribution probability of each security event context sequence as the score of the topic; Step 4.5: According to the five-stage attack model, the attack is divided into five stages: initial intrusion, establishing a foothold, lateral movement, penetration / obstruction, and post-penetration / post-obstruction. And a score is assigned to each attack stage according to the importance of each stage and the severity of the consequences caused; when an event can be classified into multiple stages, the stage with the highest score is selected; combining expert knowledge and experience, according to the relevance of events of different task categories and keywords in the dataset to different attack stages, the events are assigned to one of the stages to obtain a maliciousness rating table for the dataset. Step 4.
6. According to the theme-event distribution Select the top top_n events with the highest probabilities from each theme-event distribution; if the event sequence contains the above events, match these events with the maliciousness rating table and add the scores of the corresponding levels. Multiple events of the same level are only scored once; calculate the score s of the event sequence on each theme by accumulating the scores k , and the sum of the scores of each theme is used to evaluate the relevance between the theme and the malicious attack; then, normalize the scores of each theme to obtain the weight ω of the event sequence for each theme k , and the calculation formula is as follows: Step 4.7: Calculate the maliciousness score p for each sample event sequence mal , if it exceeds the preset threshold p0, then determine that the event sequence is a malicious attack; the calculation formula is: where θ k represents the probability that the event sequence belongs to the k-th topic in the event sequence-topic distribution θ in step 4.4; Step 4.8: Finally, according to the malicious attack determination results of the sample event sequences, the maliciousness of the entire clustering is estimated; if all sample event sequences are determined to be malicious attack sequences, it is considered that the entire clustering to which they belong belongs to malicious attacks; if no sample event sequence is determined to be a malicious attack sequence, it is considered that the entire clustering to which they belong does not belong to malicious attacks; if there are some sample event sequences determined to be malicious attack sequences, it means that the clustering contains both malicious attack sequences and normal behavior sequences, and security analysts need to further analyze the reasons, mark the clustering as ambiguous, and re-execute Steps 4.3 to 4.7 for all sample event sequences in the clustering.
Citation Information
Cited By
Security event association aggregation method and device, storage medium and product
CN121580060A