Intrusion alarm attack stage classification method, device, equipment, medium and product
By using active learning sampling strategies and semi-supervised learning strategies in intrusion alarm classification, the problem of sparse label data and fuzzy sample data labels is solved, and the accurate classification and interpretation of the intrusion alarm attack stage is achieved.
Patent Information
- Application Number
- CN202510051153.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively classify and interpret attack phases and behaviors in intrusion alerts, especially when labeling data is sparse and sample data labels are blurred.
Active learning sampling strategy and semi-supervised learning strategy are used to generate unlabeled intrusion alerts, and the language model and classifier in the field of network security are trained by updating the training set to form a classification model in the network attack stage.
In the case of sparse labeling data and blurred sample data labels, the attack phase classification of intrusion alerts is effectively carried out, which improves the judgment and response capabilities of potential attacks.
Smart Images

Figure CN120074866A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, and particularly to a method, device, equipment, medium and product for classifying the attack stages of intrusion alarms. Background Art
[0002] With the rapid development of information and communication technology, the Internet of Things and Industry 5.0, modern network attack behaviors have become more diverse, organized, and the attack means are more concealed and complex. Traditional network defense models such as firewalls and rule-based intrusion detection systems can no longer cope with diverse and complex network unknown threats. With the continuous expansion of the network scale and the increasing complexity of the threat environment, the challenges faced by the network security analysis center are also intensifying. The network security analysis center processes tens of thousands of intrusion alarms every day. Merely relying on traditional expert analysis methods, analysts are often overwhelmed by a large number of alarms, making it extremely difficult to effectively extract key information from intrusion alarms. At the same time, there are also significant differences in the sources of network intrusion alarms, resulting in the generated alarms often lacking clear context information and guiding significance, thereby affecting the judgment and response capabilities for potential attacks. This information ambiguity is the direct manifestation that intrusion alarms cannot be directly applied, seriously restricting the requirements of analysts in terms of timeliness and accuracy. How to use a small amount of labeled data to train and extract the features of unlabeled data to achieve an accurate mapping of the network attack stages of intrusion alarm data is an urgent problem to be solved. Summary of the Invention
[0003] This application provides a method, device, equipment, medium and product for classifying the attack stages of intrusion alarms, which solves the problems of scarce labeled data and fuzzy sample data labels in the task of classifying the attack stages of intrusion alarms.
[0004] This application provides a method for classifying the attack stages of intrusion alarms, including:
[0005] Obtain the intrusion alarm to be classified;
[0006] Input the intrusion alarm to be classified into a pre-trained network attack stage classification model to obtain a classification result, where the classification result includes the network attack stage and network attack behavior corresponding to the intrusion alarm to be classified;
[0007] Among them, the training process of the network attack stage classification model includes:
[0008] Generate labels for unlabeled intrusion alarms based on a preset active learning sampling strategy and semi-supervised learning strategy to obtain newly labeled intrusion alarms, and update the training set according to the newly labeled intrusion alarms;
[0009] Train the language model and classifier in the network security field according to the updated training set;
[0010] Repeat the above steps to obtain the network attack phase classification model.
[0011] As an embodiment, generating labels for unlabeled intrusion alerts based on a preset active learning sampling strategy and a semi-supervised learning strategy to obtain newly labeled intrusion alerts, and updating the training set according to the newly labeled intrusion alerts includes:
[0012] Selecting first target intrusion alerts that meet the active learning sampling strategy from unlabeled intrusion alerts based on the preset active learning sampling strategy for manual labeling to obtain first newly labeled intrusion alerts;
[0013] Selecting second target intrusion alerts with a confidence level reaching a preset threshold from unlabeled intrusion alerts based on the preset semi-supervised learning strategy to obtain second newly labeled intrusion alerts;
[0014] Adding the newly labeled intrusion alerts composed of the first newly labeled intrusion alerts and the second newly labeled intrusion alerts to the training set composed of labeled intrusion alerts to update the training set.
[0015] As an embodiment, selecting first target intrusion alerts that meet the active learning sampling strategy from unlabeled intrusion alerts based on the preset active learning sampling strategy includes:
[0016] Obtaining an uncertainty metric value according to the standard deviation of the classification probabilities output by the network security domain language model and the classifier during the dropout training process;
[0017] Determining an instance correlation value based on the last self-attention layer of the network security domain language model, where the instance correlation value is used to characterize the correlation between different words in the labeled intrusion alerts;
[0018] Determining a sampling objective function for the active learning sampling strategy according to the uncertainty metric value and the instance correlation value;
[0019] Selecting intrusion alerts that meet the sampling objective function from unlabeled intrusion alerts as the first target intrusion alerts.
[0020] As an embodiment, selecting second target intrusion alerts with a confidence level reaching a preset threshold from unlabeled intrusion alerts based on the preset semi-supervised learning strategy to obtain second newly labeled intrusion alerts includes:
[0021] Training a semi-supervised model based on the labeled intrusion alerts;
[0022] Predicting unlabeled intrusion alerts based on the semi-supervised model to obtain the prediction probabilities of each unlabeled intrusion alert;
[0023] Select intrusion alerts with confidence reaching a preset threshold from the unlabeled intrusion alerts according to the predicted probability as the second target intrusion alerts;
[0024] Pseudo-label the second target intrusion alerts according to the predicted probability to obtain the second newly labeled intrusion alerts.
[0025] As an embodiment, the training of the language model and classifier in the network security field according to the updated training set includes:
[0026] Train the pre-trained language model and pre-trained classifier according to the updated training set to obtain the parameters corresponding to the pre-trained language model and the pre-trained classifier respectively;
[0027] Based on transfer learning, transfer the parameters corresponding to the pre-trained language model and the pre-trained classifier to the language model and classifier in the network security field to obtain the parameters corresponding to the language model and classifier in the network security field respectively;
[0028] Adjust the parameters of the language model in the network security field according to the updated training set to complete the training of the language model and classifier in the network security field.
[0029] As an embodiment, the network attack phase includes a reconnaissance and scanning phase, a vulnerability exploitation phase, a persistence phase, and an attack execution phase. The network attack behaviors include host discovery, service discovery, vulnerability discovery, and information discovery corresponding to the reconnaissance and scanning phase; also include privilege escalation, brute-force credential access, attack using publicly facing application vulnerabilities, attack using remote service vulnerabilities, and arbitrary code execution corresponding to the vulnerability exploitation phase; also include defense evasion and command and control corresponding to the persistence phase; also include endpoint denial-of-service attack, network denial-of-service attack, service stop, data theft, and data delivery corresponding to the attack execution phase.
[0030] This application also provides an apparatus for classifying the attack phase of intrusion alerts. The apparatus includes:
[0031] A training module, configured to generate labels for unlabeled intrusion alerts based on a preset active learning sampling strategy and semi-supervised learning strategy to obtain newly labeled intrusion alerts, update the training set according to the newly labeled intrusion alerts; train a language model and classifier in the network security field according to the updated training set; repeat the above steps to obtain a network attack phase classification model;
[0032] An acquisition module, configured to acquire intrusion alerts to be classified;
[0033] A classification module, configured to input the intrusion alert to be classified into a pre-trained network attack stage classification model to obtain a classification result, where the classification result includes the network attack stage and network attack behavior corresponding to the intrusion alert to be classified.
[0034] This application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the intrusion alert attack stage classification method as described in any one of the above is implemented.
[0035] This application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the intrusion alert attack stage classification method as described in any one of the above is implemented.
[0036] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the intrusion alert attack stage classification method as described in any one of the above is implemented.
[0037] The intrusion alert attack stage classification method, device, equipment, medium, and product provided by this application can generate labels for unlabeled intrusion alerts through an active learning sampling strategy and a semi-supervised learning strategy, and can effectively classify the attack stage in the case of few labeled samples and fuzzy sample data labels in the training set. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 is one of the flow diagrams of the intrusion alert attack stage classification method provided by this application.
[0040] Figure 2 is the second flow diagram of the intrusion alert attack stage classification method provided by this application.
[0041] Figure 3 is the training flow diagram of the language model in the field of network security under the active learning sampling strategy provided by this application.
[0042] Figure 4 is the transfer learning training flow diagram of the language model in the field of network security provided by this application.
[0043] Figures 5a - 5cThey are respectively the comparison charts of the network attack stages output by the present application and the existing model under three evaluation metrics: training loss value, F1 value, and accuracy rate.
[0044] Figures 6a - 6d They are respectively the comparison charts of the network attack behaviors output by the present application and the existing model under four evaluation metrics: training loss value, F1 value, accuracy rate, and top3 accuracy rate.
[0045] Figure 7 It is a schematic structural diagram of the intrusion alert attack stage classification method device provided by the present application.
[0046] Figure 8 It is a schematic structural diagram of the electronic device provided by the present application. Specific implementation manners
[0047] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0048] It should be noted that all actions of obtaining signals, information, or data in the present application are carried out on the premise of complying with the corresponding data protection regulations and policies of the location and with the authorization given by the owner of the corresponding device.
[0049] Figure 1 It is one of the flow schematic diagrams of the intrusion alert attack stage classification method provided by the present application. As Figure 1 shown, the present application provides an intrusion alert attack stage classification method, including step S100-step S200.
[0050] Step S100, obtain the intrusion alert to be classified. The intrusion alert to be classified refers to the alert that needs to be processed by the network security analysis center.
[0051] Step S200, input the intrusion alert to be classified into the pre-trained network attack stage classification model to obtain a classification result, where the classification result includes the network attack stage and network attack behavior corresponding to the intrusion alert to be classified.
[0052] The network attack stage classification model is used to extract features and represent features for the intrusion alarms to be classified, that is, to convert the text sequence corresponding to the intrusion alarms to be classified into a high-dimensional semantic representation, which contains rich semantic features and context information; then, abstract and non-linearly transform the feature representation, output a probability distribution corresponding to the number of classification labels, and finally make a classification decision based on this distribution, so as to classify the intrusion alarms to be classified into the corresponding network attack stages and network attack behaviors. For example, after feature abstraction and transformation, the network attack stage classification model will finally generate a vector, and each element of the vector represents the probability of each category. This probability value can be used to represent the possibility that the intrusion alarm to be classified belongs to a specific network attack stage. For instance, if there are three attack stages (such as "reconnaissance stage", "attack stage" and "post-exploitation stage"), then the model will output a vector containing three values, respectively representing the probability that the alarm belongs to each stage. Finally, the model will determine the classification result of the alarm according to the label with the highest probability. The network attack stage classification model can achieve the task of explaining the intrusion alarms to be classified, enabling the security analysis center to intuitively perceive the attack intention of each alarm.
[0053] Among them, the training process of the network attack stage classification model includes step S010 - step S030.
[0054] Step S010, generate labels for the unlabeled intrusion alarms based on a preset active learning sampling strategy and semi-supervised learning strategy, obtain newly labeled intrusion alarms, and update the training set according to the newly labeled intrusion alarms.
[0055] The active learning sampling strategy specifically refers to selecting intrusion alarms containing important information from the unlabeled intrusion alarm sample data based on sampling factors, and the sampling factors are determined based on important information. Correspondingly, generating labels for the unlabeled intrusion alarms based on the preset active learning sampling strategy means selecting intrusion alarms based on the active learning sampling strategy, and after manual annotation of the selected intrusion alarms by network security experts, generating labels, where the labels are used to indicate the network attack stage and network attack behavior to which the intrusion alarms belong.
[0056] Important information usually refers to features or data points that can help the model quickly learn and improve performance, such as boundary samples, uncertain samples, etc.
[0057] Boundary samples refer to samples close to the decision boundary, that is, samples that the model is currently uncertain about their classification. The information they contain is very crucial for improving the decision boundary of the model. Therefore, selecting these samples for annotation can significantly improve the accuracy of the model.
[0058] For a model, uncertain samples refer to those samples with relatively low classification confidence of the model. Uncertain samples are very helpful for optimizing the model during the training process, and active learning strategies usually prefer to select these samples for annotation.
[0059] Semi-supervised learning strategy is a learning strategy that combines labeled data and unlabeled data, and can still improve the learning performance of the model in the case of a lack of a large amount of labeled data.
[0060] This application combines an active learning sampling strategy and a semi-supervised learning strategy to construct an active semi-supervised intrusion alert label generation algorithm, combining the advantages of active learning and semi-supervised learning.
[0061] The initial data of the training set consists of labeled intrusion alerts, and the training set will be updated during each training process.
[0062] Step S020: Train the language model and classifier in the field of network security according to the updated training set. In this step, samples can be directly selected from the updated training set to train the language model and classifier in the field of network security, or the idea of transfer learning can be used to train the pre-trained language model corresponding to the language model in the field of network security and the pre-trained classifier corresponding to the classifier with the updated training set, and transfer the trained parameters to the language model and classifier in the field of network security to speed up the training.
[0063] Step S030: Repeat the above steps to obtain the network attack stage classification model. Repeat Step S010 and Step S020 until the preset number of iterations is reached, or the language model and classifier in the field of network security converge, and use the combination of the language model and classifier in the field of network security as the network attack stage classification model.
[0064] It can be understood that this application generates labels for unlabeled intrusion alerts through an active learning sampling strategy and a semi-supervised learning strategy, obtains newly labeled intrusion alerts and expands the training set, and can effectively classify the attack stage in the case of few labeled samples and fuzzy sample data labels in the training set, solving the problems of scarce labeled data and fuzzy sample data labels in the intrusion alert attack stage classification task.
[0065] Based on the above embodiments, as an optional embodiment, the network attack phase includes a reconnaissance and scanning phase, a vulnerability exploitation phase, a persistence access phase, and an attack execution phase. The network attack behaviors include host discovery, service discovery, vulnerability discovery, and information discovery corresponding to the reconnaissance and scanning phase; it also includes privilege escalation, brute-force credential access, attacking using publicly facing application vulnerabilities, attacking using remote service vulnerabilities, and arbitrary code execution corresponding to the vulnerability exploitation phase; it further includes defense evasion and command and control corresponding to the persistence access phase; and it also includes endpoint denial-of-service attack, network denial-of-service attack, service stop, data theft, and data delivery corresponding to the attack execution phase.
[0066] The present application pre-defines a set of action intention states from two perspectives, macroscopic and microscopic. The macroscopic perspective describes what the attacker ultimately attempts to attack (intention), and the microscopic perspective describes how the expected goal is achieved (specific technology), which explains the specific nature and expected intention of the attack behavior corresponding to each intrusion alert. The macroscopic phases of network attacks are divided into four phases: reconnaissance and scanning, vulnerability exploitation, persistence access, and attack execution. The microscopic phase contains 16 specific attack behaviors, as shown in Table 1.
[0067] Table 1 Definition Table of Macroscopic and Microscopic Phases of Network Attacks
[0068]
[0069] The network attack phase classification model provided by the present application is used to interpret intrusion alerts with scarce and ambiguous labels as clear and direct network attack phases and network attack behaviors, describing the intrusion alerts from two perspectives, so that the security analysis center can intuitively perceive the attack intention of each alert.
[0070] It can be understood that the present application defines the action intention state of the attack behavior represented by each intrusion alert from two perspectives, which can intuitively and concisely interpret the intrusion alerts, enabling the security analysis center to intuitively perceive the attack intention of each alert.
[0071] Figure 2 is the second flow schematic diagram of the intrusion alert attack phase classification method provided by the present application. As Figure 2 shown, based on the above embodiments, as an optional embodiment, the unlabeled intrusion alerts are labeled based on a preset active learning sampling strategy and a semi-supervised learning strategy to obtain newly labeled intrusion alerts, and the training set is updated according to the newly labeled intrusion alerts, including steps S011 - S013.
[0072] Step S011: Select the first target intrusion alerts that meet the active learning sampling strategy from the unlabeled intrusion alerts based on the preset active learning sampling strategy for manual annotation to obtain the first newly labeled intrusion alerts. Specifically, select the intrusion alerts containing important information from the unlabeled intrusion alert sample data as the first target intrusion alerts based on the sampling factors, and then have cybersecurity experts manually annotate the selected intrusion alerts to generate the labels corresponding to the first target intrusion alerts. The first target intrusion alerts after generating the labels are used as the first newly labeled intrusion alerts.
[0073] Step S012: Select the second target intrusion alerts with confidence reaching the preset threshold from the unlabeled intrusion alerts based on the preset semi-supervised learning strategy to obtain the second newly labeled intrusion alerts.
[0074] Step S013: Add the newly labeled intrusion alerts composed of the first newly labeled intrusion alerts and the second newly labeled intrusion alerts to the training set composed of the labeled intrusion alerts to update the training set.
[0075] Let the labeled intrusion alert dataset, i.e., the training set, be denoted as L, the unlabeled intrusion alert dataset be denoted as U, and the model of the current training round be denoted as θ. Use the active learning sampling strategy to select samples from U and then have experts annotate them to obtain the first newly labeled intrusion alerts D 1 , and at the same time remove the data from U. Use the semi-supervised learning strategy to select samples from U and then perform pseudo-labeling to obtain the second newly labeled intrusion alerts D 2 , and at the same time remove the data from U. Combine the alerts labeled by the active learning sampling strategy and the alerts labeled by the semi-supervised learning strategy to update the training set L ′ =D 1 +D 2 , L = L + L ′ .
[0076] It should be noted that Figure 2 the SecureBERT model in
[0077] It can be understood that compared with all unknown intrusion alarms, the labeled intrusion alarms are extremely limited. The performance of a model trained only with the labeled intrusion alarms is poor. Therefore, this application combines an active learning sampling strategy and a semi-supervised learning strategy to obtain an active semi-supervised intrusion alarm label generation algorithm, combining the advantages of active learning and semi-supervised learning. Active learning can intelligently select samples for labeling for limited labeled data, thereby reducing the labeling cost, but there is a problem of over-reliance on manual and professional domain knowledge. Semi-supervised learning can use unlabeled data to increase the training samples and improve the generalization ability of the model. However, semi-supervised learning has a bias of insufficient professional knowledge, resulting in a risk of model overfitting. Combining the two can more comprehensively utilize the limited labeled data and a large amount of unlabeled data to train the model, solve the difference between the labeled data set and the unknown data, thereby improving the training process of the model, accelerating the model training, and improving the model classification accuracy.
[0078] Based on the above embodiments, as an optional embodiment, selecting the first target intrusion alarms that meet the active learning sampling strategy from the unlabeled intrusion alarms based on the preset active learning sampling strategy includes steps S0111 - S0114.
[0079] Step S0111, obtaining an uncertainty metric value according to the standard deviation of the classification probabilities output by the network security domain language model and the classifier during the dropout training process.
[0080] Step S0112, determining an instance correlation value based on the last self-attention layer of the network security domain language model, where the instance correlation value is used to characterize the relevance between different words in the labeled intrusion alarms.
[0081] Step S0113, determining a sampling objective function of the active learning sampling strategy according to the uncertainty metric value and the instance correlation value.
[0082] Step S0114, selecting the intrusion alarms that meet the sampling objective function from the unlabeled intrusion alarms as the first target intrusion alarms.
[0083] Figure 3 is a schematic diagram of the training process of the network security domain language model under the active learning sampling strategy provided by this application. As Figure 3 shown, the embodiments of this application will describe the active learning sampling strategy in detail in combination with the preferred embodiments of the network security domain language model.
[0084] In the field of network security, the language model preferably adopts the SecureBERT model. The training process of the SecureBERT model starts from a small amount of labeled data and a large amount of unlabeled data. The SecureBERT model includes multiple encoding layers. This application uses the last encoding layer to mine the instance correlation between instances, and combines the uncertainty measurement to form a utility function. The utility function is used to select the intrusion alerts containing important information from the unlabeled intrusion alert data pool for manual annotation based on two perspectives of uncertainty and correlation, and then convert them into labeled intrusion alerts, which are added to the training process to accelerate the model training and improve the model accuracy.
[0085] It is challenging for deep learning models to calculate their uncertainty because the models themselves are not interpretable, and the output class probabilities are not a reliable method for evaluating uncertainty. This application uses the idea of eliminating neurons during training to improve generalization. By randomly discarding a certain percentage of neurons in multiple iterations, the standard deviation of the classification probabilities in multiple dropout iterations is calculated as the uncertainty measure. The SecureBERT model is f nn , the data instance is x, the number of Dropout iterations is T, and the configuration of the i-th Dropout is d i , the calculation formula for the average value of the output probabilities of the model in T iterations:
[0086]
[0087] Among them, p is the average value of the output probability.
[0088] The uncertainty measure is defined as the distribution dispersion of the prediction probabilities of all models, and the calculation formula is as follows:
[0089]
[0090] Among them, R uncertain (X) is the uncertainty measure value.
[0091] For a neural network model, an uncertain prediction will have a higher distribution dispersion value (i.e., uncertainty measure), indicating that different parts of the neural network have conflicting activations for a given input. The uncertainty measure can help select samples that are challenging for the model for annotation, thereby improving the model performance.
[0092] For the correlation metric, this application explores the correlation of instances in the last self-attention layer of the SecureBERT model. Instance correlation is used to measure the interaction between words in each text. By calculating the average of the variances of the multi-head attention, the degree of variation between different heads can be obtained, which reflects the correlation between different words in the sentence. In the self-attention mechanism, each word takes into account the importance of other words, which is reflected in the attention weight matrix. By calculating the sum of the attention of each word, the importance of each word in the entire sentence can be understood. At the same time, by calculating the average of the variances of the multi-head attention, the degree of variation between different heads can be obtained, reflecting the correlation between different words in the sentence. Therefore, the correlation within the sentence can be well represented through this process. Strongly correlated instances indicate that the model is more confident in them. The specific calculation strategy for the correlation metric is as follows: Let AM i be the attention matrix of head i, which reflects the influence of all words in the text on the updated word. The internal correlation is measured by jointly considering the interaction between words in each text, and its calculation formula is:
[0093]
[0094] where AM i represents the attention matrix of the i-th head, with a size of d vocab ×d vocab , a ij represents the element in the i-th row and j-th column of the attention matrix AM i , and n represents the number of words in the text.
[0095] Sum the rows of AM i to calculate the sum of the attention of each word, obtaining S i . The intra-instance correlation R intra (X) is the average of the variances of all heads, and the calculation formula is as follows:
[0096]
[0097] where S i represents the sum of the attention of each word in the attention matrix of the i-th head, s j is the j-th component of the vector S i , representing the total attention value of the j-th word in the entire sentence, represents the average of the sums of the attention of all heads, used to characterize the average importance of each word in the sentence, R intra(X) represents the instance correlation value, which is used to measure the correlation between words in a sentence. h is the number of heads in the multi-head attention mechanism. In the multi-head self-attention mechanism, each head independently calculates the attention weights, and then aggregates the outputs of all heads. In the correlation measurement, the formula calculates the correlation of the attention matrix AM i for each head, and then takes the average of all heads to measure the overall correlation of the instance.
[0098] The calculation formula of the sampling objective function is as follows:
[0099]
[0100] where X sel represents the sampling objective function.
[0101] From the perspective of uncertainty measurement, an uncertain prediction will have a higher variance value of the prediction probability (i.e., uncertainty measurement), indicating that different parts of the neural network are ambiguous about the given input. From the perspective of instance correlation measurement, the internal correlation of an instance refers to the logical connection and coherence between various parts of the instance. If the internal correlation of an instance is high, it means that the meaning expressed by the instance is clear, and the logical relationship between various parts is tight. The more confident the model is about the correlation of the instance, the clearer the model thinks the structure and meaning of the instance are, and the logical relationship between various parts is correct.
[0102] It can be understood that this application constitutes an active learning sampling strategy by fusing uncertainty and instance correlation to actively select instances (i.e., unlabeled intrusion alerts containing important information), no longer relying on a single strategy, but considering both the uncertainty and correlation of sampling at the same time, and selecting the most representative samples from unlabeled data to augment the training set, effectively enhancing the interpretability of label-scarce and fuzzy intrusion alerts, achieving an accurate mapping between fuzzy intrusion alerts and accurate network attack stages, and solving the problems of scarce labeled data and fuzzy information in the intrusion alert attack stage classification task.
[0103] Based on the above embodiments, as an optional embodiment, the step of selecting second target intrusion alerts with a confidence level reaching a preset threshold from unlabeled intrusion alerts based on a preset semi-supervised learning strategy to obtain second newly labeled intrusion alerts includes steps S0121 - S0124.
[0104] Step S0121: Train a semi-supervised model based on the labeled intrusion alerts.
[0105] Step S0122: Predict the unlabeled intrusion alerts based on the semi-supervised model to obtain the prediction probability of each unlabeled intrusion alert.
[0106] Step S0123: Select intrusion alerts with a confidence level reaching a preset threshold from the unlabeled intrusion alerts according to the predicted probability as the second target intrusion alerts.
[0107] Step S0124: Pseudo-label the second target intrusion alerts according to the predicted probability to obtain the second newly labeled intrusion alerts.
[0108] The semi-supervised learning strategy is mainly used to expand the training data. Specifically, the semi-supervised learning process includes the following steps: First, use a training set composed of labeled data, which includes initially labeled instances, data labeled in each iteration of active learning, and pseudo-labels generated through the semi-supervised learning process. Next, apply the model to predict the unlabeled sample set and extract the predicted probability of each sample. Then, select a subset of samples with high confidence based on the predicted probability. Finally, pseudo-label the samples in this subset according to the prediction results of the model, and then update the training set.
[0109] It can be understood that in this application, the training set is expanded through the semi-supervised learning strategy, and accurate interpretation of intrusion alerts can be effectively achieved even when the number of labeled samples is insufficient.
[0110] Figure 4 is a schematic diagram of the transfer learning training process of the language model in the field of network security provided by this application. As Figure 4 shown, on the basis of the above embodiment, as an optional embodiment, the training of the network security field language model and the classifier according to the updated training set includes steps S021 - S022.
[0111] Step S021: Train the pre-trained language model and the pre-trained classifier according to the updated training set to obtain the parameters corresponding to the pre-trained language model and the pre-trained classifier respectively.
[0112] Step S022: Based on transfer learning, transfer the parameters corresponding to the pre-trained language model and the pre-trained classifier to the network security field language model and the classifier respectively to obtain the parameters corresponding to the network security field language model and the classifier respectively.
[0113] Step S023: Adjust the parameters of the network security field language model according to the updated training set to complete the training of the network security field language model and the classifier.
[0114] The attack phase classification model based on transfer learning proposed in this application is based on the combination of SecureBERT and DenseNN, aiming to use the pre-trained SecureBERT model for feature extraction and representation learning, and implement further processing of the feature representation and execution of classification tasks through the DenseNN layer.
[0115] The SecureBERT model is a language model in the field of network security based on the BERT model. In the architecture of this application, SecureBERT is responsible for converting the input text sequence into a high-dimensional semantic representation, which contains rich semantic features and context information.
[0116] Based on the semantic representation output by the SecureBERT model, DenseNN (fully connected neural network) acts as a classifier, receiving the semantic representation from the SecureBERT model as input, further abstracting and non-linearly transforming the features through a multi-layer fully connected structure, and finally outputting a prediction result that conforms to the number of classification labels.
[0117] Next, the technical effects of the preferred embodiments of this application will be elaborated in detail in combination with simulation experiment data. This application was compared with four other baseline models under different data ratios. The four baseline models are shown in Table 2.
[0118] Table 2 Definition Table of Baseline Models
[0119] model description Base BERT Base BERT model secureBERT Base secureBERT model Active - learning - only SecureBERT secureBERT model with active - learning strategy Semi - supervised - learning - only SecureBERT secureBERT model with semi - supervised - learning strategy
[0120] This application selects intrusion alert data from the security competitions CPTC and CCDC to approximate the fuzzy intrusion alerts captured in a real network environment. The experiment uses 10-fold cross-validation. Each experiment randomly divides the dataset to avoid repetition. The training data and test data are divided in a ratio of 4:1. Active learning is simulated by reserving a part of the labeled dataset for expert annotation. This application uses top1 accuracy, top3 accuracy, F1 value, and Loss value as evaluation indicators. The top1 and top3 accuracies represent the percentage of the correct results predicted by the model in the total samples and the percentage of the number of samples in which one of the top three results predicted by the model is correct in the total samples, respectively. The F1 value makes an overall evaluation based on the accuracy rate and recall rate.
[0121] Figures 5a - 5c They are respectively the comparison charts of the network attack phases output by this application and the existing models under three evaluation indicators of training loss value, F1 value, and accuracy. Figures 6a - 6d They are respectively the comparison charts of the network attack behaviors output by this application and the existing models under four evaluation indicators of training loss value, F1 value, accuracy, and top3 accuracy. The horizontal axis in the figure is the number of training rounds, such as Figures 5a - 6dAs shown in the figure, SecureBERT has better performance than the basic BERT model; only semi-supervised learning has large fluctuations because it uses the pseudo-labeling mode, and the predicted data may not improve performance and has weak robustness; only active learning quickly reaches an inflection point and then the accuracy basically remains unchanged because the remaining sample information content is low and has little impact on the model. The method proposed in this application can quickly find effective samples to enter the training set, save the number of labeled samples and training costs, and more comprehensively evaluate the value of samples. From the perspective of performance indicators, this method is superior to active and semi-supervised learning in the entire iterative process in terms of macro and micro attack stage explanations, and has a higher accuracy rate.
[0122] In summary, this application proposes a method for classifying the attack stage of intrusion alarms based on active and semi-supervised learning, aiming to address the problems of scarce labeled data and fuzzy information in the task of classifying the attack stage of intrusion alarms. It can efficiently extract the potential attack stage information in intrusion alarms at a low cost. Even when the number of labeled samples is insufficient, it can effectively achieve accurate interpretation of alarms. This application also proposes an active learning sampling strategy that combines uncertainty and instance relevance, selects the most representative samples from unlabeled data to expand the training set, effectively enhances the ability to interpret labels-scarce and fuzzy intrusion alarms, and realizes the accurate mapping between fuzzy intrusion alarms and accurate network attack stages. Through the verification of CPTC and CCDC experimental data, the results show that this application performs excellently in different baseline models, the influence of the amount of labeled data, and the comparison of different active learning algorithms. In the case of scarce labeled intrusion alarms, it can significantly improve the interpretation effect of intrusion alarms, effectively improve the classification accuracy and training efficiency of the model.
[0123] The device for the method of classifying the attack stage of intrusion alarms provided by this application will be described below. The device for the method of classifying the attack stage of intrusion alarms described below can be mutually corresponding and referred to the method of classifying the attack stage of intrusion alarms described above.
[0124] Figure 7 It is a schematic structural diagram of the device for the method of classifying the attack stage of intrusion alarms provided by this application. As Figure 7 shown, this application also provides a device for the method of classifying the attack stage of intrusion alarms, including the following modules:
[0125] A training module 710, configured to generate labels for unlabeled intrusion alarms based on a preset active learning sampling strategy and semi-supervised learning strategy to obtain newly labeled intrusion alarms, update the training set according to the newly labeled intrusion alarms; train a language model and a classifier in the field of network security according to the updated training set; repeat the above steps to obtain a network attack stage classification model;
[0126] An acquisition module 720, configured to acquire intrusion alarms to be classified;
[0127] A classification module 730, configured to input the intrusion alert to be classified into a pre-trained network attack phase classification model to obtain a classification result, where the classification result includes a network attack phase and a network attack behavior corresponding to the intrusion alert to be classified.
[0128] As an embodiment, the training module 710 is further configured to:
[0129] Select first target intrusion alerts that meet the active learning sampling strategy from unlabeled intrusion alerts based on a preset active learning sampling strategy for manual annotation and obtain first newly labeled intrusion alerts;
[0130] Select second target intrusion alerts with a confidence level reaching a preset threshold from unlabeled intrusion alerts based on a preset semi-supervised learning strategy to obtain second newly labeled intrusion alerts;
[0131] Add the newly labeled intrusion alerts composed of the first newly labeled intrusion alerts and the second newly labeled intrusion alerts to a training set composed of labeled intrusion alerts to update the training set.
[0132] As an embodiment, the training module 710 is further configured to:
[0133] Obtain an uncertainty metric value according to the standard deviation of the classification probabilities output by the network security domain language model and the classifier during the dropout training process;
[0134] Determine an instance correlation value based on the last self-attention layer of the network security domain language model, where the instance correlation value is used to characterize the correlation between different words in the labeled intrusion alerts;
[0135] Determine a sampling objective function of the active learning sampling strategy according to the uncertainty metric value and the instance correlation value;
[0136] Select intrusion alerts that meet the sampling objective function from unlabeled intrusion alerts as the first target intrusion alerts.
[0137] As an embodiment, the training module 710 is further configured to:
[0138] Train a semi-supervised model according to the labeled intrusion alerts;
[0139] Predict unlabeled intrusion alerts based on the semi-supervised model to obtain the prediction probabilities of each unlabeled intrusion alert;
[0140] Select intrusion alerts with a confidence level reaching a preset threshold from unlabeled intrusion alerts according to the prediction probabilities as the second target intrusion alerts;
[0141] Pseudo-label the second target intrusion alert according to the predicted probability to obtain the second newly labeled intrusion alert.
[0142] As an embodiment, the training module 710 is further configured to:
[0143] Train the pre-trained language model and the pre-trained classifier according to the updated training set to obtain the parameters corresponding to the pre-trained language model and the pre-trained classifier respectively;
[0144] Based on transfer learning, transfer the parameters corresponding to the pre-trained language model and the pre-trained classifier to the language model and the classifier in the field of network security to obtain the parameters corresponding to the language model and the classifier in the field of network security respectively;
[0145] Adjust the parameters of the language model in the field of network security according to the updated training set to complete the training of the language model and the classifier in the field of network security.
[0146] As an embodiment, the network attack phase includes a reconnaissance and scanning phase, a vulnerability exploitation phase, a persistent access phase, and an attack execution phase. The network attack behaviors include host discovery, service discovery, vulnerability discovery, and information discovery corresponding to the reconnaissance and scanning phase; it also includes privilege escalation, brute-force credential access, attacking using publicly facing application vulnerabilities, attacking using remote service vulnerabilities, and arbitrary code execution corresponding to the vulnerability exploitation phase; it also includes defense evasion and command and control corresponding to the persistent access phase; it also includes endpoint denial-of-service attack, network denial-of-service attack, service stop, data theft, and data delivery corresponding to the attack execution phase.
[0147] The intrusion alert attack phase classification method device provided by this application is used to implement the intrusion alert attack phase classification method provided in any of the above embodiments, and has the corresponding technical effects of the intrusion alert attack phase classification method, which will not be elaborated here.
[0148] Figure 8 Illustrates a schematic physical structure diagram of an electronic device, such as Figure 8As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logical instructions in the memory 830 to execute the intrusion alert attack phase classification method, which includes: obtaining the intrusion alert to be classified; inputting the intrusion alert to be classified into a pre-trained network attack phase classification model to obtain a classification result, where the classification result includes the network attack phase and network attack behavior corresponding to the intrusion alert to be classified; among them, the training process of the network attack phase classification model includes: generating labels for unlabeled intrusion alerts based on a preset active learning sampling strategy and semi-supervised learning strategy to obtain newly labeled intrusion alerts, and updating the training set according to the newly labeled intrusion alerts; training a network security domain language model and a classifier according to the updated training set; repeating the above steps to obtain the network attack phase classification model.
[0149] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0150] On the other hand, the present application also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the intrusion alert attack stage classification method provided by the above-mentioned various methods. The method includes: obtaining an intrusion alert to be classified; inputting the intrusion alert to be classified into a pre-trained network attack stage classification model to obtain a classification result, where the classification result includes a network attack stage and a network attack behavior corresponding to the intrusion alert to be classified; wherein, the training process of the network attack stage classification model includes: generating labels for unlabeled intrusion alerts based on a preset active learning sampling strategy and a semi-supervised learning strategy to obtain newly labeled intrusion alerts, and updating the training set according to the newly labeled intrusion alerts; training a network security domain language model and a classifier according to the updated training set; repeating the above steps to obtain the network attack stage classification model.
[0151] On the other hand, the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the intrusion alert attack stage classification method provided by the above-mentioned various methods. The method includes: obtaining an intrusion alert to be classified; inputting the intrusion alert to be classified into a pre-trained network attack stage classification model to obtain a classification result, where the classification result includes a network attack stage and a network attack behavior corresponding to the intrusion alert to be classified; wherein, the training process of the network attack stage classification model includes: generating labels for unlabeled intrusion alerts based on a preset active learning sampling strategy and a semi-supervised learning strategy to obtain newly labeled intrusion alerts, and updating the training set according to the newly labeled intrusion alerts; training a network security domain language model and a classifier according to the updated training set; repeating the above steps to obtain the network attack stage classification model.
[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for classifying attack stages of intrusion alarms, characterized in that: include: Get intrusion alerts to be classified; Inputting the intrusion alarm to be classified into a pre-trained network attack stage classification model to obtain a classification result, wherein the classification result includes the network attack stage and network attack behavior corresponding to the intrusion alarm to be classified; The training process of the network attack stage classification model includes: Generate labels for unlabeled intrusion alarms based on a preset active learning sampling strategy and a semi-supervised learning strategy to obtain new labeled intrusion alarms, and update the training set according to the new labeled intrusion alarms; Train the network security domain language model and classifier based on the updated training set; Repeat the above steps to obtain the network attack stage classification model.
2. The intrusion alarm attack stage classification method according to claim 1 is characterized in that: The method generates labels for unlabeled intrusion alarms based on a preset active learning sampling strategy and a semi-supervised learning strategy to obtain new labeled intrusion alarms, and updates the training set according to the new labeled intrusion alarms, including: Based on a preset active learning sampling strategy, a first target intrusion alarm that meets the active learning sampling strategy is selected from unlabeled intrusion alarms for manual labeling to obtain a first newly labeled intrusion alarm; Based on a preset semi-supervised learning strategy, a second target intrusion alarm whose confidence reaches a preset threshold is selected from the unlabeled intrusion alarms to obtain a second newly labeled intrusion alarm; The newly labeled intrusion alarm consisting of the first newly labeled intrusion alarm and the second newly labeled intrusion alarm is added to a training set consisting of labeled intrusion alarms to update the training set.
3. The intrusion alarm attack stage classification method according to claim 2 is characterized in that: The method of selecting a first target intrusion alarm that complies with the active learning sampling strategy from unlabeled intrusion alarms based on a preset active learning sampling strategy includes: Obtaining an uncertainty measurement value according to the network security domain language model and a standard deviation of the classification probability output by the classifier during the discard training process; Determining an instance relevance value based on a last self-attention layer of the network security domain language model, wherein the instance relevance value is used to characterize the relevance between different words in the annotated intrusion alert; Determining a sampling objective function of an active learning sampling strategy according to the uncertainty metric value and the instance relevance value; An intrusion alarm that meets the sampling objective function is selected from unlabeled intrusion alarms as the first target intrusion alarm.
4. The intrusion alarm attack stage classification method according to claim 2 is characterized in that: The method of selecting a second target intrusion alarm whose confidence reaches a preset threshold from unlabeled intrusion alarms based on a preset semi-supervised learning strategy to obtain a second newly labeled intrusion alarm includes: A semi-supervised model is trained based on labeled intrusion alerts; Predicting unlabeled intrusion alarms based on the semi-supervised model to obtain a prediction probability of each unlabeled intrusion alarm; Selecting, from the unlabeled intrusion alarms, an intrusion alarm whose confidence reaches a preset threshold value as the second target intrusion alarm according to the predicted probability; The second target intrusion alarm is pseudo-labeled according to the predicted probability to obtain the second newly labeled intrusion alarm.
5. The intrusion alarm attack stage classification method according to any one of claims 1 to 4, characterized in that: The training of the network security domain language model and classifier according to the updated training set includes: Training the pre-trained language model and the pre-trained classifier according to the updated training set to obtain parameters corresponding to the pre-trained language model and the pre-trained classifier respectively; Based on transfer learning, the parameters corresponding to the pre-trained language model and the pre-trained classifier are transferred to the network security domain language model and the classifier to obtain the parameters corresponding to the network security domain language model and the classifier; The parameters of the network security domain language model are adjusted according to the updated training set to complete the training of the network security domain language model and the classifier.
6. The intrusion alarm attack stage classification method according to any one of claims 1 to 4, characterized in that: The network attack stages include a reconnaissance and scanning stage, a vulnerability exploitation stage, a maintenance access stage and an attack implementation stage. The network attack behaviors include host discovery, service discovery, vulnerability discovery and information discovery corresponding to the reconnaissance and scanning stage; and also include privilege escalation, brute force credential access, attacks using public application vulnerabilities, attacks using remote service vulnerabilities and arbitrary code execution corresponding to the vulnerability exploitation stage; and also include defense avoidance and command and control corresponding to the maintenance access stage; and also include endpoint denial of service attacks, network denial of service attacks, service shutdown, data theft and data delivery corresponding to the attack implementation stage.
7. A method and device for classifying attack stages of intrusion alarms, characterized in that: include: A training module is used to generate labels for unlabeled intrusion alarms based on a preset active learning sampling strategy and a semi-supervised learning strategy to obtain new labeled intrusion alarms, and to update a training set according to the new labeled intrusion alarms; to train a network security domain language model and a classifier according to the updated training set; and to repeat the above steps to obtain a network attack stage classification model; An acquisition module, used for acquiring intrusion alarms to be classified; The classification module is used to input the intrusion alarm to be classified into a pre-trained network attack stage classification model to obtain a classification result, which includes the network attack stage and network attack behavior corresponding to the intrusion alarm to be classified.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the intrusion alarm attack stage classification method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the intrusion alarm attack stage classification method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the intrusion alarm attack stage classification method according to any one of claims 1 to 6 is implemented.