An industrial intrusion detection method based on BILSTM-CRF
By adopting the BILSTM-CRF model in industrial control network intrusion detection, combined with the advantages of Bi-LSTM and CRF, the problem of low classification accuracy of traditional single deep learning algorithms is solved, and higher classification accuracy and better industrial information security performance are achieved.
Patent Information
- Application Number
- CN202210856731.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-07-21
AI Technical Summary
Traditional single deep learning algorithms have low classification accuracy in industrial control network intrusion detection, making it difficult to effectively identify multiple attack categories.
The BILSTM-CRF model combined with bidirectional long and short-term memory network (Bi-LSTM) and conditional random field (CRF) is used to improve the classification accuracy by performing timing processing of network intrusion data and estimating the probability of tag sequences.
Experimental results on the CICIDS2017 dataset show that the accuracy and F1 value of the BILSTM-CRF model have been significantly improved, and the classification accuracy reaches 99.81%, which is far higher than other models, meeting the real-time intrusion detection needs of industrial control systems.
Smart Images

Figure CN115396143B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an industrial control network intrusion detection method, and in particular to an industrial intrusion detection method based on BILSTM-CRF. Background Art
[0002] Internet services have become an indispensable part of business transactions and personal life. With the increasing reliance on network services, the confidentiality and integrity of critical information are increasingly compromised by remote intrusions. Enterprises are forced to strengthen their networks against malicious activities and cyber threats. Therefore, network systems must use one or more security tools, such as firewalls, antivirus software or intrusion detection systems, to protect important data and services from hackers or intruders.
[0003] Because firewalls cannot protect against intrusion attempts on open ports required for network services, relying on firewall systems alone is not enough to protect a company's network from all types of network attacks. Therefore, it is usually necessary to install an intrusion detection system (IDS) to supplement the firewall. The intrusion detection system collects information from the network or computer system and analyzes this information to detect system intrusions.
[0004] Network intrusion detection has always been a problem that needs to be solved urgently in network security. Traditional intrusion detection algorithms have problems such as low detection rate and high false alarm rate. Summary of the invention
[0005] The purpose of the present invention is to provide an industrial intrusion detection method based on BILSTM-CRF. The present invention solves the problem of low classification accuracy of the traditional single deep learning algorithm classification model of industrial control network intrusion, and uses BILSTM-CRF combined with BILSTM and CRF to improve the classification accuracy.
[0006] The objective of the present invention is achieved through the following technical solutions:
[0007] A BILSTM-CRF based industrial intrusion detection method, the method comprising the following steps:
[0008] Step 1: Preprocess the intrusion detection dataset and divide it into a training set and a test set;
[0009] Step 2: Preprocess the raw data traffic collected in the network;
[0010] Step 3: Build a BILSTM-CRF network model, pass the training set into the BILSTM-CRF network model, and the output layer outputs the predicted values of samples in different categories;
[0011] Step 4: The MLP classifier can be used to implement the classification task of intrusion detection; the processed data passes through two fully connected layers and one Dropout layer, and is finally output through the Softmax function.
[0012] The invention discloses an industrial intrusion detection method based on BILSTM-CRF, which uses a bidirectional (forward, backward) LSTM coupled with CRF as a configuration of a neural network of the output layer.
[0013] The method is based on BILSTM-CRF industrial intrusion detection. The input of each time step of the method is a data, which is converted into a data vector by a data embedding layer. In addition to the data vector, the input layer can also selectively adopt additional data-based features to utilize feature representation.
[0014] The industrial intrusion detection method based on BILSTM-CRF is described, wherein the surface concatenated embedding vector and additional features are placed in the forward and backward layers of LSTM respectively; for each time step, the outputs of the forward LSTM and the backward LSTM are connected and sent to the hidden layer; the output that is fully connected to the hidden layer gives multi-class probability information corresponding to the label; finally, the probability information of the network element label on the entire data sequence is placed in the CRF layer to estimate the optimal network element label sequence on the entire text sequence.
[0015] The industrial intrusion detection method based on BILSTM-CRF, the CRF layer estimates the label of each data position, and jointly maximizes the label sequence through the Viterbi algorithm; the BILSTM layer and the CRF layer are connected to form a complete BILSTM-CRF model.
[0016] The advantages and effects of the present invention are:
[0017] Compared with other industrial abnormal intrusion detection methods, the present invention has a specific improvement in that, in response to the increasingly serious industrial information security problem, the present invention proposes a network intrusion detection model composed of a bidirectional long short-term memory network (Bi-LSTM) and a conditional random field (CRF) based on the temporal characteristics of network intrusions—the BILSTM-CRF model. Experiments were conducted using the public data set CICIDS2017 of the Canadian Institute for Cybersecurity (CIC) at the University of New Brunswick (UNB). The experimental results show that the model of the present invention has improved the precision P and F1 values of various attack categories of network intrusion detection compared to traditional machine learning models, and the final accuracy of the model is higher (99.81%), which is much higher than other models, meeting the requirements of network intrusion detection. The superiority of the performance of the model in the application of real-time intrusion detection in industrial control systems has been demonstrated. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram of the steps of the method of the present invention;
[0019] Figure 2 This is the schematic diagram of the BILSTM network structure;
[0020] Figure 3 This is the schematic diagram of the BILSTM-CRF network model. DETAILED DESCRIPTION
[0021] The present invention will be described in detail below with reference to the embodiments shown in the accompanying drawings.
[0022] Recurrent neural networks (RNNs) can process temporal information based on historical information, which gives them a natural advantage when dealing with temporal data. However, due to the gradient vanishing and gradient exploding problems, RNNs will have problems such as long-term dependency. LSTM is a memory unit specifically designed to solve long-term dependency problems. It is mainly composed of three types of gate structures, which control the proportion of information to be forgotten and stored in the cell state. For the task of extracting temporal data, if we access past and future data information at a given time, we can use more levels of information to make better predictions. However, LSTM cannot encode information from back to front. BILSTM can encode information in both forward and reverse order, which solves the problem of not being able to encode from front to back. The specific structure is as follows Figure 2 . The main advantage of CRF is that it can incorporate rich, overlapping features. It incorporates dependencies between observations and aims to solve the problem of long-range dependencies. In CRF, features based on the rich relationship between input and output vectors can be easily incorporated. Therefore, it does not require modeling effort on observations, which are fixed at test time. In addition, the conditional probability of the label sequence can depend on arbitrary, non-independent features of the observation sequence without forcing the model to consider these feature dependencies.
[0023] BILSTM-CRF layer
[0024] The input to the BiLSTM layer is a sequence of representations (vectors) of the input data tokens, expressed as The output of the BiLSTM layer is a sequence of hidden states for each input data vector as follows:
[0025]
[0026] The final output vector is the forward input to each hidden state and the backward hidden state The combined vector is as follows:
[0027]
[0028] CRF Layer
[0029] CRF is the most commonly used method for control structure prediction. Its basic idea is to use a series of potential functions to approximate the conditional probability of the output label sequence given an input sequence.
[10] Formally, we take the above hidden state sequence h = As our input to the CRF layer, its output is our final predicted label sequence y = , In the set of all possible labels, we will It is represented as a set of all possible label sequences. Then, given the input hidden state sequence, the conditional probability of the output sequence is derived as:
[0030]
[0031] Where W and b are two weight matrices, indicating that we extract the given label pair To train the CRF layer, we use the classic maximum conditional likelihood estimation to train our model. The final log-likelihood of the weight matrix is:
[0032]
[0033] BILSTM-CRF model
[0034] The concatenated embedding vector and additional features of the surface are placed in the forward and backward layers of the LSTM respectively. For each time step, the outputs of the forward LSTM and the backward LSTM are connected and sent to the hidden layer. The output that is fully connected to the hidden layer gives the multi-class probability information corresponding to the label. Finally, the probability information of the network element label on the entire data sequence is put into the CRF layer to estimate the optimal network element label sequence on the entire text sequence. The CRF layer estimates the label of each data position and jointly maximizes the label sequence through the Viterbi algorithm. Connect the BILSTM layer and the CRF layer to form a complete BILSTM-CRF model. The specific structure is as follows: Figure 2 .
[0035] MLP Classifier
[0036] The MLP classifier is a fully connected neural network, which mainly consists of two parts: the input layer and the output layer. The BILSTM-CRF hybrid network is used to extract advanced features from the sample data. Only a simple classifier, MLP, is needed to achieve the classification task of intrusion detection. First, the processed data passes through two fully connected layers and a Dropout layer, and finally outputs through the Softmax function. The output result of the BILSTM-CRF hybrid network is P, and the MLP score calculation formula is as follows:
[0037]
[0038] A BISTM-CRF-based industrial intrusion detection method, the specific steps are as follows:
[0039] Step 1: Preprocess the intrusion detection dataset and divide it into a training set and a test set;
[0040] Step 2: Preprocess the raw data traffic collected in the network;
[0041] Step 3: Build a BILSTM-CRF network model, pass the training set into the BILSTM-CRF network model, and the output layer outputs the predicted values of samples in different categories;
[0042] Step 4: The MLP classifier can be used to achieve the classification task of intrusion detection. The processed data passes through two fully connected layers and a Dropout layer, and finally outputs through the Softmax function;
[0043] Embodiment 1:
[0044] The operating system used in the experiment is Windows 10 64-bit system, the CPU is Intel I7-10700, the running memory is 8G, and the deep learning framework is tensorflow2.7.
[0045] Experimental data source
[0046] The public dataset of the Canadian Institute for Cyber Security (CIC) at the University of New Brunswick (UNB) is used for the intrusion detection (IDS) task. The CICIDS2017 dataset contains 3.1 million flow records. After removing incomplete records, there are still about 2.9 million flow data. The Canadian Institute for Cyber Security at the University of New Brunswick also provides complete packet captures of CICIDS2017, but these data are not used in this paper. The 84 continuous features are min-max normalized, and the dataset sample labels are One-Hot encoded. The description of the dataset after preprocessing is shown in Table 1.
[0047] Table 1 Dataset description
[0048]
[0049] Evaluation indicators
[0050] In order to objectively evaluate the performance of the D-IDS model, the accuracy (ACC), precision (P), detection rate (R) and comprehensive evaluation index (F1-Measure, F) commonly used in the field of intrusion detection are used to evaluate the classification results.
[0051]
[0052] Among them, TP represents the number of correctly identified attack categories, FN represents the number of missed attacks, i.e., the number of incorrectly identified attack categories, FP represents the number of false positives, i.e., the number of incorrectly identified normal categories, and TN represents the number of correctly identified normal categories.
[0053] Parameter settings
[0054] In order to obtain better experimental results, the BILSTM-CRF model is used for one iteration with an appropriate number of input model samples BATCH-SIZE=500, a dropout rate Dropout-rate=0.5, EPOCHS=20 for all sample training, an initial learning rate LR=0.001, and an optimization function using Adam.
[0055] Comparison Algorithms
[0056] In the intrusion detection model, the commonly used machine learning algorithms include artificial neural network (ANN), decision tree (DT), K-nearest neighbor (KNN), naive Bayes (NB), random forest (RF), support vector machine (SVM), k-means clustering algorithm (K-MEANS), expectation-maximization algorithm (EM), and self-organizing map (SOM). These commonly used algorithms are used to build intrusion detection models and test the CICIDS2017 data set. The comparison of the detection results of different machine learning models is shown in Figure 2.
[0057] Table 2 Comparison of detection of different models
[0058] .
Claims
1. A BILSTM-CRF based industrial intrusion detection method, characterized in that: The method comprises the following steps: Step 1: Preprocess the intrusion detection dataset and divide it into a training set and a test set; Step 2: Preprocess the raw data traffic collected in the network; Step 3: Build a BILSTM-CRF network model, pass the training set into the BILSTM-CRF network model, and the output layer outputs the predicted values of samples in different categories; Step 4: Use the MLP classifier to implement the classification task of intrusion detection; the processed data passes through two fully connected layers and one Dropout layer, and finally outputs through the Softmax function; The input of the BILSTM layer is a vector sequence of input data labels, represented as The output of the BILSTM layer is the hidden state sequence of each input data vector as shown in formula (1) and formula (2): ; ; The final output vector is the forward hidden state sequence and the backward hidden state sequence The combined vector is shown in formula (3): ; CRF Layer The CRF layer takes a hidden state sequence given an input sequence. As input to the CRF layer, its output is the final predicted label sequence , In the set of all possible labels; Represented as a set of all possible label sequences; then given an input hidden state sequence Under the condition of , the conditional probability of the output sequence is derived as shown in formula (4): ; Where W and b are two weight matrices, For a given label pair, the final log-likelihood formula of the weight matrix is shown in formula (5): ; BILSTM-CRF model The concatenated embedding vector and additional features of the surface are placed in the forward and backward layers of the LSTM respectively; for each time step, the outputs of the forward LSTM and the backward LSTM are connected and sent to the hidden layer; the output that is fully connected to the hidden layer gives the multi-class probability information corresponding to the label; finally, the probability information of the network element label on the entire data sequence is put into the CRF layer to estimate the optimal network element label sequence on the entire text sequence; the CRF layer estimates the label of each data position and jointly maximizes the label sequence through the Viterbi algorithm; the BILSTM layer and the CRF layer are connected to form a complete BILSTM-CRF model; MLP Classifier The MLP classifier is a fully connected neural network, which mainly consists of two parts: the input layer and the output layer. The sample data is feature extracted through the BILSTM-CRF hybrid network, and only the MLP classifier is needed to achieve the classification task of intrusion detection. First, the processed data passes through two fully connected layers and a Dropout layer, and finally is output through the Softmax function.