Industrial control network flow intrusion detection method combining deep learning and multi-kernel learning

By combining deep learning and multi-core learning methods, high-dimensional features of industrial control network traffic are extracted and MDKELM models are generated, which solves the problem of insufficient accuracy and efficiency in industrial control network intrusion detection, and achieves intrusion detection effects with high accuracy and real-time response.

CN120200794APending Publication Date: 2025-06-24SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332042.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Traditional machine learning algorithms have problems of insufficient accuracy, efficiency and robustness in industrial-controlled network intrusion detection, especially when dealing with high-dimensional nonlinear features and real-time requirements.

Method used

Using a combination of deep learning and multi-core learning, the intermediate layer expression is extracted through deep neural network (DNN) to form a kernel matrix, and the multi-core learning method is used to combine the kernel matrix to generate an MDKELM model, and deploy it on edge devices for intrusion detection.

Benefits of technology

It improves the accuracy and efficiency of traffic intrusion detection of industrial control networks, achieves a multi-classification accuracy of 90.09%, and performs excellently in real-time response and model parameters, adapting to the real-time and robustness requirements of industrial control services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200794A_ABST
    Figure CN120200794A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of industrial network security, and particularly relates to an industrial control network traffic intrusion detection method combining deep learning and multi-kernel learning, which comprises the following steps of: firstly, acquiring an industrial control traffic data sample and performing feature cleaning, and then establishing a multilayer deep neural network DNN model; and the optimal parameter configuration of the model is obtained through iterative training. Then extracting a middle layer representation result of the DNN as a kernel matrix of mapping; the kernel matrixes obtain respective weights through a multi-kernel learning MKL process, and linear combination is carried out. And finally, the combined kernel matrix replaces the original shallow kernel function of the kernel extreme learning machine KELM, and a multi-depth kernel extreme learning machine MDKELM model for detecting the traffic sample type is formed. The model is deployed at the edge of a device and a network, external access traffic is detected, whether the external access traffic belongs to normal traffic is judged, and if not, early warning is given out in time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention patent belongs to the field of industrial network security, and specifically relates to an industrial control network traffic intrusion detection method combining deep learning and multi-core learning. Background Art

[0002] In an industrial control network traffic intrusion detection system (IDS), the intrusion detection model after effective training is the core component. It can discover and prevent potential intrusion behaviors by monitoring network and system activities in real time, thereby improving system security, reducing the impact of security incidents, and enhancing the efficiency of security management. The performance of the model directly affects the performance of the entire IDS. With the rapid development of computers and the emergence and popularization of intelligent technologies, many machine learning methods have been applied to industrial control network intrusion detection and achieved certain results. However, traditional machine learning algorithms have certain limitations, and the simple IDS designed thereby lacks in terms of accuracy, efficiency, and robustness. In addition, the network traffic information in industrial control operations has the characteristics of high-dimensional non-linearity, which places relatively strict requirements on the mapping ability of machine learning models, as well as the real-time requirements of industrial control operations themselves. These reasons result in most IDSs based on traditional machine learning algorithms being difficult to detect various attacks or abnormal accesses suffered by devices quickly and accurately in actual production applications, and the false alarms generated by them may also cause losses to production operations. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides an industrial control network traffic intrusion detection method combining deep learning and multi-core learning, including the following steps:

[0004] Step 1: Obtain industrial control traffic data from industrial production processes or traffic logs and perform feature cleaning.

[0005] Step 2: Build a DNN model and train it using the dataset obtained in Step 1.

[0006] Step 3: Extract the intermediate layer expressions of the DNN network to form a kernel matrix.

[0007] Step 4: Combine the kernel matrix using the multi-core learning method to generate an MDKELM model, deploy the model on edge devices, and perform intrusion detection.

[0008] Further, obtain network traffic samples accessing the target device within a certain period of time, or obtain samples from traffic logs, numericalize the character-type information therein, and then assign labels to all samples, that is, whether the sample belongs to normal traffic or a certain type of attack traffic, and integrate them into a neural network dataset.

[0009] Further, obtain data samples and perform feature cleaning, specifically as follows:

[0010] Step 1.1: Obtain network traffic samples accessing the target device within a certain period of time, or obtain samples from traffic logs, numericalize the character features therein, and then assign labels to all samples, that is, whether the sample belongs to normal traffic or a certain type of attack traffic, and integrate them into a neural network training dataset. The term "feature" is a machine learning term, which refers to various attribute information in traffic data samples in this article, such as source IP address, target IP address, industrial control protocol type, etc. Generally speaking, each column of data in the dataset represents a feature, and each row of data represents a sample.

[0011] Step 1.2: First, according to the mRMR (Max-Relevance and Min-Redundancy) principle, calculate the mutual information measure between each feature and the label, as shown in formula (1); this process can calculate the degree of association between various features in the collected traffic samples; this measure can be used to measure the importance of the features of the samples:

[0012]

[0013] where J mRMR (x k ) is the feature contribution degree of x k , y is the target vector, S is the feature space, containing all features; x j represents the j-th feature among other features in S except x k , j refers to any one except k, and I(y:x k ) represents the mutual information measure, and the calculation formula is as follows:

[0014]

[0015] where p(x k ,y) represents the joint probability distribution of x k and y, and p(x k ) represents the independent probability of x k .

[0016] Step 1.3: Use J mRMR (x k ) to measure the classification contribution degree of the k-th feature, and rank all features; then, intercept an appropriate proportion of features as input data to participate in subsequent model training.

[0017] Furthermore, the training process of the DNN model is as follows:

[0018] Step 2.1: Establish a feedforward DNN model, whose structure is 1 input layer, 4 Linear layers, that is, fully connected intermediate layers, and 1 output layer, including a Softmax basic classifier;

[0019] Step 2.2: Input the data with features cleaned into the input layer of the DNN for tensorization, and convert the data size to n*m*1, where n is the number of samples and m is the number of features;

[0020] Step 2.3: After tensorization, the data is passed into the middle layer for high-dimensional mapping. The filter sizes of the middle layer are 32, 32, 64, and 64 respectively, that is, the data is mapped into 32, 32, 64, and 64 dimensions respectively in the middle layer;

[0021] Step 2.4: Calculate the loss according to the Softmax classification calculation result of the input layer and the cross-entropy loss function, and perform backpropagation to train the model.

[0022] Furthermore, use the DNN model trained in Step 2 to generate deep kernels, specifically as follows:

[0023] The middle layer Φ of the DNN can be used to construct a kernel generation function; regard Φ as a feature vector in the high-dimensional space, and then use the kernel function to map points in the low-dimensional space to the high-dimensional space; the definition of the kernel function generated by the neural network is:

[0024] K r (x1,x2) = <Φ r (x1), Φ r (x2)>

[0025] The mapping matrix of each hidden layer in the DNN participates in the multiple kernel learning (MKL) as the kernel matrix; each middle layer representation in the DNN model is also extracted to generate the corresponding deep kernel; after the optimal linear combination through the multiple kernel learning method, the expressions of these deep kernels are as follows:

[0026]

[0027] where L is the number of hidden layers, and μ r represents the weight coefficient of each kernel, which is obtained through the MKL process.

[0028] Furthermore, use the multiple kernel learning method to combine the deep kernels generated in the previous step to generate the MDKELM model. The function definition of ELM is as follows:

[0029] h i (x) = g(ω i x + b i )

[0030]

[0031] Among them, ω represents the weight vector between the input layer and the hidden layer, β represents the weight vector between the hidden layer and the output layer, and b represents the bias of the hidden nodes. g(x) represents the activation function, f(x) represents the output result, and H represents the output matrix of the hidden layer.

[0032] The basic kernel function of KELM is defined as follows:

[0033]

[0034] The calculation formula of KELM can be expressed in the following form:

[0035]

[0036] Among them, Ω r is the r-th kernel function or kernel matrix, f(x) is the output result, T is the target matrix (label), and I is the identity matrix;

[0037] Applying the kernel matrix extracted in step 3 here, the calculation expression of MDKELM can be obtained:

[0038]

[0039] Among them, μ r is calculated according to the FHeuristic multi-kernel learning method, as follows:

[0040]

[0041] This is a heuristic MKL method that assigns weights according to the proximity of the individual kernel to the ideal kernel; this method is particularly effective because it does not require solving any optimization problems except for the final training of the base classifier; in the above definition, A(K1, K2) represents the alignment method between two kernel matrices, and yy T represents the ideal kernel; as shown in the following formula:

[0042]

[0043] Finally, the output f(x) of MDKELM is the probability that the input traffic x belongs to different labels (normal or a certain attack type), and the label with the highest probability is selected as the final classification result of the MDKELM model.

[0044] The advantages of the present invention are:

[0045] Aiming at the industry pain points such as the difficulty in analyzing high-dimensional non-linear characteristics of industrial control network traffic, high real-time detection latency, and high false alarm rate of traditional models, the present invention innovatively integrates deep neural network (DNN) and multi-kernel learning (MKL) technologies, and proposes a deep multi-kernel intrusion detection method suitable for industrial control services. By constructing a deep kernel matrix through feature mapping in the middle layer of DNN and combining a heuristic multi-kernel dynamic weighting mechanism, multi-level feature fusion of industrial control protocol traffic is realized, effectively identifying complex threats such as APT attacks and protocol camouflage. Experimental detection is carried out on public data sets such as NSL-KDD, and the multi-classification accuracy rate is increased to 90.09%; a real-time response at the 35ms level is achieved on edge devices, and the number of model parameters is compressed to less than 1MB. The present invention designs a lightweight deep neural network structure, selects a suitable multi-kernel learning method, reduces the detection time cost of the model, and enables the model to adapt to large-scale data by streamlining the kernel, so as to meet the real-time and robustness requirements of industrial control services. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a schematic flow chart of the present invention;

[0047] Figure 2 is a schematic module structure diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0049] Refer to Figure 1-2 , the present invention provides an industrial control network traffic intrusion detection method combining deep learning and multi-kernel learning, including the following steps:

[0050] Step 1: Obtain industrial control traffic data from industrial production processes or traffic logs and perform feature cleaning

[0051] MDKELM obtains network traffic samples accessing the target device within a certain period of time, or obtains samples from traffic logs, numerically converts the character-type information therein, and then assigns labels to all samples (that is, whether the sample belongs to normal traffic or a certain type of attack traffic), and integrates them into a neural network training dataset. It is divided into a training set, a validation set, and a test set according to the ratio of 0.6:0.2:0.2. The features are various attribute information in the traffic data samples. Since the effect of neural network feature mapping is closely related to the feature quality. Therefore, it is necessary to select valuable features before model training. Information-Theoretic Feature Selection (ITFS) is one of the most effective tools, which can eliminate these irrelevant and redundant features while retaining relevant features. Compared with the Pearson correlation coefficient, ITFS can capture non-linear relationships and is not affected by the feature scale. Compared with feature filtering, ITFS not only evaluates the correlation between features and the target variable, but also considers the relationships between features, so as to be able to discover the effect of feature combinations. More importantly, it does not require additional model training, greatly reducing the cost of feature engineering.

[0052] Step 1.1: First, according to the "Max-Relevance and Min-Redundancy" (mRMR) principle, as shown in formula (1), calculate the mutual information measure between each feature and the label. In practical applications, this process can calculate the degree of association between various attribute information (such as IP address and access time, etc.) in the traffic samples collected by the model. This measure can be used to measure the importance of the k-th feature of the sample, where k refers to any one:

[0053]

[0054] where, J mRMR (x k ) is the feature contribution degree of x k , y is the target vector, S is the feature space, including all features; x j represents the j-th feature among other features in S except x k , where j refers to any one except k, and I(y:x k ) represents the mutual information measure, and the calculation formula is as follows:

[0055]

[0056] where, p(x k ,y) represents the joint probability distribution of x k and y, and p(x k ) represents the independent probability of x k .

[0057] Step 1.2: Use J mRMR(x k ) to measure the classification contribution of the k-th feature and rank all features. Then, intercept an appropriate proportion of features as input data to participate in subsequent model training.

[0058] Step 2: Build a DNN model and train it using the dataset obtained in Step 1.

[0059] The principle of a deep neural network is to provide a complex representation of input data through the stacking of non-linear mappings, that is, to achieve high-dimensional mapping of data, namely to transform the traffic attribute information implemented in Step 1.2 into high-dimensional information that can be understood by the neural network. The calculations performed by this model can be mathematically described as a non-linear function:

[0060] Φ(x) = g(Wx + b)

[0061] where W is the weight matrix and b is the bias vector. For a general deep feedforward neural network, its mapping function can be expressed as follows:

[0062]

[0063] where Φ1 maps the input to the first hidden layer, and each operation represents a non-linear transformation from hidden layer r - 1 to hidden layer r.

[0064] Step 2.1: Build a feedforward DNN model with a structure of 1 input layer, 4 Linear layers (fully connected intermediate layers), and 1 output layer (including a Softmax basic classifier).

[0065] Step 2.2: Input the feature-cleaned data into the input layer of the DNN, perform tensorization, and convert the data size to n * m * 1, where n is the number of samples and m is the number of features.

[0066] Step 2.3: After tensorization, the data is passed into the intermediate layer for high-dimensional mapping. The filter sizes of the intermediate layer are 32, 32, 64, 64 respectively, that is, the data is mapped into 32, 32, 64, 64 dimensions respectively in the intermediate layer.

[0067] Step 2.4: Calculate the loss according to the Softmax classification calculation result of the input layer and the cross-entropy loss function, and perform backpropagation to train the model.

[0068] Step 3: Extract the intermediate layer expression of the DNN network to form a kernel matrix.

[0069] Inspired by deep learning, the kernel can be connected to a series of non-linear functions to achieve a highly flexible new kernel function. The principle of using a neural network to generate a kernel is to utilize the neural network to learn representations by increasing the complexity of the feature hierarchy, and extract the feature maps generated by the intermediate layers of the neural network model as kernel matrices for calculation. The deep mapping kernel based on a deep neural network has an explicit expression of non-linear mapping and can be extended from training data to test data in an analytical solution manner. In addition, the deep kernel can effectively extract the features contained in the data and remove the noise and redundant information in the data.

[0070] When processing data, DNN automatically maps the data features to a specified number of dimensions, i.e., the number of channels. Therefore, the intermediate layer Φ representation of DNN can be used to construct the generated kernel function. Specifically, Φ can be regarded as a feature vector in a high-dimensional space, and then the kernel function is used to map the points in the low-dimensional space to the high-dimensional space. The definition of the neural network-generated kernel function is:

[0071] K r (x1,x2) = <Φ r (x1), Φ r (x2) >

[0072] The mapping matrix of each hidden layer in DNN can participate in multi-kernel learning (MKL) as a kernel matrix and has a positive effect. Therefore, we also extract the representation of each intermediate layer in the DNN model to generate the corresponding deep kernel. After the optimal linear combination through the multi-kernel learning method, the expressions of these deep kernels are as follows:

[0073]

[0074] where L is the number of hidden layers, and μ r represents the weight coefficient of each kernel, obtained through the MKL process. What the kernel (or kernel matrix) is responsible for in this method is to transform the feature information extracted by DNN into a complex non-linear mapping in a higher dimension.

[0075] Step 4: Use the multi-kernel learning method to combine the kernel matrices to generate the MDKELM model.

[0076] The basic kernel function definition of KELM is shown as follows:

[0077]

[0078] The calculation formula of KELM can be expressed in the following form:

[0079]

[0080] where Ω r$K_r$ is the $r$-th kernel function or kernel matrix, $f(x)$ is the output result, $T$ is the target matrix (label), and $I$ is the identity matrix.

[0081] Applying the kernel matrix extracted in Step 3 here, the calculation expression of MDKELM can be obtained:

[0082]

[0083] where $\mu$ r is calculated according to the FHeuristic multi-kernel learning method as follows:

[0084]

[0085] This is a heuristic MKL method that assigns weights based on the proximity of individual kernels to the ideal kernel. This method is particularly effective because it does not require solving any optimization problems except for the final training of the base classifier. In the above definition, $A(K_1, K_2)$ represents the alignment method between two kernel matrices, and $y^y$ T represents the ideal kernel. As shown in the following formula.

[0086]

[0087] Finally, the output $f(x)$ of MDKELM is the probability that the input traffic $x$ belongs to different labels (normal or a certain attack type), and the label with the highest probability is selected as the final classification result of the MDKELM model.

[0088] Step 5: Deploy the MDKELM model on the edge device and perform intrusion detection. The process is the same as the training process, but the intermediate DNN no longer requires iterative training. The data to be tested can be input into the MDKELM model after feature cleaning to output the final classification result.

[0089] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for detecting intrusion in industrial control network traffic by combining deep learning and multi-core learning, characterized in that: The following steps are included: Step 1: Obtain industrial control flow data from industrial production processes or flow logs and perform feature cleaning to obtain a data set; Step 2: Build a multi-layer deep neural network DNN model and train it using the data set obtained; Step 3: Extract the intermediate layer expression of the DNN network to form a kernel matrix; Step 4: Use the multi-kernel learning method to combine the kernel matrix and generate the multi-deep kernel extreme learning machine MDKELM model to perform intrusion detection on industrial control traffic data.

2. The industrial control network traffic intrusion detection method combining deep learning and multi-core learning as claimed in claim 1, characterized in that: Obtaining data samples and performing feature cleaning specifically includes the following steps: Step 1.1: Obtain network traffic samples that access the target device within a certain period of time, or obtain samples from traffic logs, and digitize the character features therein, then assign labels to all samples, that is, whether the samples belong to normal traffic or a certain attack traffic, and integrate them into a neural network data set, which is divided into a training set, a validation set, and a test set in a ratio of 0.6:0.2:0.

2. The features are various attribute information in the traffic data samples; Step 1.2: Perform feature cleaning through feature selection method based on information theory. First, calculate the contribution of each feature according to the mRMR (Max-Relevance and Min-Redundancy) principle, as shown in the following formula. This process calculates the degree of correlation between the features in the collected traffic samples, which is called mutual information metric. This metric can be used to measure the importance of the kth feature of the sample, where k refers to any one: Among them, J mRMR (x k ) is x k The feature contribution of x, y is the target vector, S is the feature space, including all features; j Indicates that S is divided by x k The jth feature among the features other than k, j refers to any one except k, I(y:x k ) represents the mutual information metric, and the calculation formula is as follows: Among them, p(x k ,y) represents x k The joint probability distribution of and y, p(x k ) represents x k The independent probability of Step 1.3: Using J mRMR (x k ) is used to measure the classification contribution of the kth feature and rank all features; then, features with high contribution are intercepted in appropriate proportions as input data for subsequent model training.

3. The industrial control network traffic intrusion detection method combining deep learning and multi-core learning as claimed in claim 1, characterized in that: The process of building a DNN model and training it is as follows: Step 2.1: Establish a feedforward DNN model with a structure of 1 input layer, 4 Linear layers, i.e. fully connected intermediate layers, and 1 output layer, including a Softmax basic classifier; Step 2.2: Input the feature-cleaned data into the input layer of DNN for tensorization, and convert the data size to n*m*1, where n is the number of samples and m is the number of features; Step 2.3: After tensorization, the data is passed to the middle layer for high-dimensional mapping. The filter sizes of the middle layer are 32, 32, 64, and 64, respectively. That is, the data is mapped into 32, 32, 64, and 64 dimensions in the middle layer respectively; Step 2.4: Calculate the result of Softmax classification according to the input layer, calculate the loss according to the cross entropy loss function, and perform back propagation to train the model.

4. The industrial control network traffic intrusion detection method combining deep learning and multi-core learning as claimed in claim 1, characterized in that: Use the DNN model trained in step 2 to generate the deep kernel as follows: The middle layer Φ of DNN can be used to construct a generative kernel function; Φ is regarded as a feature vector in a high-dimensional space, and then a kernel function is used to map points in a low-dimensional space to a high-dimensional space. The neural network generates a kernel function K r The definition is: K r (x1,x2)=<Φ r (x1),Φ r (x2)> The mapping matrix of each hidden layer in DNN participates in multi-core learning MKL as a kernel matrix; after the optimal linear combination is performed through the multi-core learning method, the final deep kernel K D The expression is as follows: Where L is the number of hidden layers, μ r Represents the weight coefficient of each kernel, obtained through the MKL process.

5. The industrial control network traffic intrusion detection method combining deep learning and multi-core learning as claimed in claim 1, characterized in that: The deep kernels generated in the previous step are combined using the multi-core learning method to generate the MDKELM model; the specific steps are as follows: The function definition of the extreme learning machine ELM is as follows: h i (x)=g(ω i x+b i ) Wherein, ω represents the weight vector between the input layer and the hidden layer, β represents the weight vector between the hidden layer and the output layer, and b represents the bias of the hidden node; g(x) represents the activation function, f(x) represents the output result, and H represents the hidden layer output matrix; The basic kernel function definition of the kernel extreme learning machine KELM is as follows: The calculation formula of KELM can be expressed as follows: Among them, f(x) is the output result, T is the target matrix (label), and I is the unit matrix; The calculation expression of MDKELM can be obtained by weighted combination of the kernel matrix (local kernel) extracted in step 3 and the basic kernel function (global kernel): Where Ω r K r (x i ,x j ), the output f(x) is the probability that the input traffic x belongs to different labels, normal or a certain attack type, and the label with the highest probability is the final classification result of the MDKELM model; It is the expression of the combination of the local kernel and the global kernel; μ r It is calculated based on the FHeuristic multi-kernel learning method, as shown below: This is a heuristic MKL method that assigns weights to individual kernels based on their proximity to the ideal kernel; in the above definition, A(K1, K2) represents the alignment between the two kernel matrices, yy T represents the ideal kernel; as shown in the following formula: In this process, MDKELM will use the training data set to calculate and obtain the appropriate μ r Parameters are the training process of the multi-core ELM classifier; after training, MDKELM can be deployed on edge devices to perform real-time detection of network traffic.