Industrial control network flow classification method for data link layer
By adopting a semi-supervised classifier based on K-Means secondary clustering and improving K nearest neighbor density peak clustering in industrial control networks, the problem of poor detection effect of imbalanced data sets is solved, and high accuracy and flexibility classification of industrial control flow is achieved.
Patent Information
- Application Number
- CN202510362098.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-01
AI Technical Summary
The existing industrial control network traffic classification method has poor detection effect when processing unbalanced data sets and high scene enclosure, resulting in low accuracy and difficulty in applying in the industrial Internet of Things.
A method of industrial control network traffic classification for data link layer is proposed, using a semi-supervised classifier based on K-Means secondary clustering and an improved K nearest neighbor density peak clustering algorithm, combining real-time cutting and collection and preprocessing technology to achieve accurate classification of industrial control traffic.
By effectively utilizing limited label data and identifying known and unknown industrial control protocols, the accuracy and flexibility of classification are improved, the dependence on data quality is reduced, and the processing capability of unbalanced data sets is enhanced.
Smart Images

Figure CN120234650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly to an industrial control network traffic classification method for the data link layer. Background Art
[0002] An industrial control network (Industrial Control System, ICS) is an automatic control system composed of the combination of computer technology and process control components. This system includes key components such as sensors, controllers, actuators, transmitters, and input / output interfaces. These components are connected through industrial communication paths and follow specific communication protocols to jointly form an industrial production or processing system with automatic control capabilities.
[0003] Network traffic classification technology is a key technology commonly used to identify and classify various data flows transmitted through a network. Network traffic classification plays a crucial role in multiple fields such as network management, security monitoring, quality of service assurance, and user behavior analysis. For example, network traffic classification can help security analysts quickly identify potential malicious activities, such as botnet communications, spyware communications, unauthorized data transmissions, etc.; security systems can distinguish normal and abnormal traffic patterns, thus effectively warning of potential security threats; enterprises and service providers can understand users' network usage habits and preferences.
[0004] Current research on network traffic classification mainly focuses on the application of machine learning algorithms, especially deep learning models, which demonstrate excellent capabilities in processing network traffic data with multi-dimensional features. Deep learning models can automatically identify and extract valuable features from large amounts of data, enabling them to master complex data patterns and internal correlations during the training process. Therefore, compared with traditional port- and load-based methods, in practical applications, these models usually perform better when dealing with large-scale data sets and can achieve more accurate classification results. In addition, some advanced machine learning algorithms can also automatically identify new attack patterns or traffic generated by unknown applications in network traffic, which is difficult to achieve with traditional methods. Compared with non-machine learning methods based on fixed rules, machine learning-based traffic classification technology can usually more effectively adapt to new traffic patterns and complex network environments.
[0005] The paper "Yao Z, Ge J, Wu Y, et al. Encrypted traffic classification based on Gaussian mixture models and Hidden Markov Models[J]. Journal of Network and Computer Applications, 2020, 166:102711." proposed a traffic classification model MGHMM based on Gaussian mixture models and Hidden Markov models to complete protocol classification and identify obfuscated traffic through experiments. This method can handle dynamic ports and encrypted traffic, improving the adaptability to complex network environments, but the computational cost is relatively high.
[0006] The paper "Wang Y, Yun X, Zhang Y, et al. A multi-scale feature attention approach to network traffic classification and its model explanation[J]. IEEE Transactions on Network and Service Management, 2022, 19(2):875 - 889." proposed using a convolutional neural network (CNN) as a building block for a deep packet analysis model, performing in-depth analysis using the raw byte sequence of individual packets to identify the generated application protocols or applications, thereby achieving accurate network traffic classification. This method uses a convolutional neural network (CNN) as a building block for a deep packet analysis model. However, the classification performance of this method highly depends on the quality of the input data and the accuracy of preprocessing, without considering the handling of imbalanced data.
[0007] In summary, existing industrial control network traffic classification methods have problems such as poor detection effect on imbalanced data sets and overly closed scenarios, resulting in low accuracy, difficulty in being applied in industrial Internet of Things, and low practicality and reference value. Summary of the Invention
[0008] The object of the present invention is to propose an industrial control network traffic classification method for the data link layer, and adopt a combined classification strategy to achieve accurate classification of industrial control traffic. (1) Considering the scarcity of industrial control traffic data, a real-time cutting and collection method for data link layer traffic data applicable to field devices of industrial control systems is designed. (2) A semi-supervised classifier based on K-Means secondary clustering is proposed, which can effectively use limited labeled data to identify known and unknown industrial control protocols, improving flexibility and accuracy in practical applications. (3) A semi-supervised combined classification method is proposed, which combines improved K-nearest neighbor density peak clustering and classifier prediction results to enhance the classification accuracy and reliability of industrial control traffic.
[0009] The technical solution of the present invention is as follows: An industrial control traffic classification method for the data link layer, including the following steps:
[0010] Step 1: Perform real-time cutting, collection and preprocessing on the industrial control traffic of the data link layer to obtain a reduced-dimensional data matrix Z;
[0011] Step 2: Construct a semi-supervised classifier based on the nearest cluster, combine limited labeled data and a large amount of unlabeled data, and adopt secondary clustering technology to improve the reliability of clustering;
[0012] Step 3: Propose a density peak clustering algorithm based on improved K-nearest neighbor, cluster the industrial control traffic data, calculate the local density and distance peak of each data point, and divide the industrial control traffic data into different clusters;
[0013] Step 4: Combine the classification results of the semi-supervised classifier based on the nearest cluster with the clustering results of Step 3 to form a combined classification strategy and obtain the final classification result.
[0014] The real-time cutting and collection in Step 1 are specifically as follows:
[0015] The real-time cutting and collection of the industrial control traffic of the data link layer are divided into a loose start stage, a truncation start stage and a recovery stage;
[0016] In the loose start stage, in order to ensure the rationality and effectiveness of traffic cutting while obtaining enough valid data, the active time threshold Tact is initially set to T1, 120≤T1≤150; check the traffic sessions obtained after splitting the industrial control traffic of the data link layer. If it is detected that the split-truncated traffic session is complete, it indicates that there is a deviation between the current Tact value and the ideal truncation time, and the Tact value is reduced; in this stage, the update of Tact follows an exponential law of decrease, and the formula is as follows:
[0017] Tact = T / 2
[0018] When Tact continuously decreases and is lower than the threshold Tthresh, it enters the truncation start phase; the update of Tact follows a linear law of decrease, and the update formula is as follows:
[0019] Tact = Tact - T2, (4 ≤ T2 ≤ 8)
[0020] After the industrial control traffic is truncated at the data link layer, when it is found that the traffic session is incomplete during the traffic session check, it immediately enters the recovery phase; the truncation start threshold Tthresh is set to twice the current Tact, and Tact is reset to 120 seconds, and then it returns to the loose start phase again.
[0021] The preprocessing in the first step includes data alignment and principal component analysis based on singular value decomposition;
[0022] The data alignment evaluates this goal by measuring the number of bits supplemented or lost due to padding or truncation operations relative to the original data of the industrial control traffic at the data link layer; the formula for the value range of the alignment length L is defined as follows:
[0023]
[0024] min(|x1|,..., |x n |) ≤ L ≤ max(|x1|,..., |x n |)
[0025] where, x i represents the i-th binary industrial control traffic;
[0026] The principal component analysis based on singular value decomposition includes data centering, singular value decomposition, principal component extraction and dimensionality reduction;
[0027] The data centering is specifically:
[0028] X c = X - μ
[0029] where, the data matrix X is an m×n matrix, where m is the number of samples and n is the number of features, and μ is the mean vector of each feature in the data matrix X;
[0030] The singular value decomposition performs singular value decomposition on the data matrix X after data centering c :
[0031] X c = U∑V T
[0032] where, U is an m×m left singular matrix, ∑ is an m×n diagonal matrix containing singular values; V T is an n×n right singular matrix;
[0033] The extraction and dimensionality reduction of the principal components are specifically as follows: The principal components are extracted from the first f columns of the right singular matrix V T of V f , where f is the number of principal components required. The data after dimensionality reduction is represented as:
[0034] Z = X c V f
[0035] where V f is the first f columns of the right singular value matrix V T , and Z is the data matrix after dimensionality reduction.
[0036] The optimal value of the alignment length L is the length value that minimizes the sum of the padding bits and the lost bits.
[0037] The semi-supervised classifier based on the nearest cluster is specifically as follows:
[0038] Step 2.1: Use the K-Means clustering algorithm to divide the data matrix after dimensionality reduction of ginger into s clusters: Randomly select s points in the dataset as the initial centers. For each data point x i , calculate its distance from all cluster centers μ j , find a cluster center that minimizes the distance, and assign the data point x i to the set C i corresponding to this cluster; For each cluster S j , calculate the coordinate mean of all points belonging to this cluster, and use the mean as the new cluster center μ j ; Minimize the sum of the squared distances of all points to their respective cluster centers, so that the data points within the same cluster are as close as possible, thereby improving the clustering effect;
[0039]
[0040]
[0041] where x i represents a data point in the dataset, μ j ∈{μ1, μ2,... μ s} represents all current cluster centers, d(x i , μ j ) represents the distance from x i to the centroid μ j ; μ j represents the new centroid of the cluster S j ; m i represents the centroid of the cluster C i , ||x j -m i|| represents the point x j to the centroid m i distance;
[0042] Step 2.2: According to the clustering results, apply the probability assignment mechanism to map the clusters created by K-Means to categories based on different industrial control protocols, and at the same time identify unknown industrial control protocols;
[0043] P(Y = y j |C i ) = n ij / n i
[0044] where P(Y = y j |C i ) represents the probability that the cluster C i is correctly mapped to the j-th category, y j represents the traffic category, n ij is the number of data points assigned to the cluster C i and whose category is j, and n is the total number of labeled data points assigned to the cluster C i ; the more the number of data points labeled as category j in the cluster C i n ij , the greater the probability that all data points in the cluster belong to category j;
[0045] Step 2.3: Based on the posterior function, the preprocessed data is divided into three main categories, namely unknown category, known category, and fuzzy category; for the cluster C i that does not contain any labeled data, it is defined as the unknown category; for the cluster containing labeled data, judge whether it belongs to the known category or the fuzzy category according to the posterior probability;
[0046] Step 2.4: Set a threshold thre, and calculate the maximum value of the posterior function for each cluster; when the maximum value of the posterior function of a certain cluster C i exceeds the threshold thre, then all data points in the cluster C i are classified into the same category, and the category of all data points in the cluster is marked as the traffic category y corresponding to the maximum value of the posterior function;
[0047]
[0048] Step 2.5: For the data whose maximum value of the posterior function does not exceed the threshold thre, use the concept of fuzzy clustering to process it; by repeating the K-Means clustering step, perform reclustering within the cluster; after completing the secondary clustering, directly perform category division according to the new clustering results and directly perform category assignment according to the corresponding criteria;
[0049] Step 2.6. After completing the secondary clustering, recalculate the clustering center points for all the generated clusters; the set of clustering centers is denoted as M = {m1, m2, m3, …, m k′}, where k′ represents the number of clusters after the secondary clustering; one industrial control protocol may correspond to one or more clustering clusters; the industrial control protocol categories are represented by Y = {y1, y2, y3, …, y q}, and for any category y i , it is represented by the centroid set M i of the sample points belonging to this category;
[0050] M i = {m j : C j ∈ y i}
[0051] Step 2.7. The classification rule of the semi-supervised classifier based on the nearest cluster is defined as follows: for a given data point x, calculate the centroid m with the minimum distance to it, find the centroid set M i to which m belongs, and assign the label of the data point x to the category j represented by the centroid set M i ;
[0052]
[0053] The specific method of the density peak clustering method improved based on K-nearest neighbors is as follows:
[0054] Step 3.1. Calculate the distance matrix D ij of the data set X using the Euclidean distance;
[0055]
[0056] where d(x i , x k ) represents the distance between x i and x k , and x ij and x kj respectively represent the values of the data points x i and x k on the j-th feature;
[0057] Step 3.2. For each point x i , calculate its k nearest neighbors in the set {x1, x2, …, x n};
[0058] Arrange the Euclidean distances related to x ij in the distance matrix D i in ascending order, and the distance ranked at the k-th position is used as the k-th nearest neighbor of x i , denoted as kNN(xi ) represents the k nearest neighbors of x i ;
[0059] Sort({d i1 , d i2 , …, d iN )
[0060] kNN(x i ) = {j ∈ X | d(x i , x k ) ≤ d(x i , NN k (x i ))}
[0061] where d(x i , x j ) represents the distance between x i and x j , NN(x i ) represents the k-th nearest neighbor of x i , and kNN(x i ) is a set containing all the points in the dataset that are within the range of the k-th nearest neighbor of x i ;
[0062] Step 3.3. Calculate the local density ρ i in the density peak clustering algorithm, and the calculation method of the local density ρ i is as follows:
[0063]
[0064] where the setting of k is based on a specific percentage of the total number N of data points, and its calculation formula is expressed as:
[0065] k = p × N
[0066] where N represents the total number of data points in the dataset, and p represents a percentage value between 0 and 1;
[0067] Step 3.4. The calculation method of the minimum distance δ i of the point x i is defined according to the following two cases; when there is at least one point x j whose density ρ j is greater than the density ρ i of the point x i , then δ i is defined as the minimum distance between the point x i and all points x j ; when there is no point whose density is greater than the point x i , then the point x i is the density peak point, and δi is defined as the maximum value of the distance between point x i and any other point in the dataset;
[0068]
[0069] Step 3.5. Consider both the local density ρ i and the minimum distance value δ i and select the points with both high local density values ρ i and large distances from other high-density points as the clustering centers; for those points not selected as clustering centers, assign them to the cluster to which their nearest point with a higher density belongs, ensuring that each point belongs to a well-defined cluster.
[0070] The specific steps of Step Four are as follows:
[0071] Step 4.2. Classify each piece of industrial control traffic data by inputting it into the semi-supervised classifier based on the nearest cluster through Step Two, and divide each piece of traffic data into finer categories;
[0072] Step 4.1. Use the density peak clustering algorithm based on the improved K-nearest neighbor (KNN-DPC) to perform preliminary clustering on the industrial control traffic dataset for the industrial control traffic data after dimensionality reduction through Step Three;
[0073] Step 4.3. Integrate the output results of each individual classifier through the combination function θ to form the final output result F(x) of the combined classifier;
[0074] F(x) = θ x∈X (f(x))
[0075] where F(x) represents the output result of the combined classifier, representing the final classification decision for the sample x, and f(x) represents the output result of an individual classifier, that is, the prediction result of each individual classifier for the sample x;
[0076] Step 4.4. Adopt the majority voting rule to summarize the prediction results of the individual classifiers to determine the final classification decision;
[0077]
[0078] where v ij is used to record the voting situation of the i-th individual classifier for the category ω j , and y xi represents the predicted category of the i-th classifier for the sample x;
[0079] Step 4.5. Select the category with the most votes from all classifiers through the majority voting rule as the final classification result of the sample;
[0080]
[0081] Among them, represents the total number of votes of category ω j which accumulates the votes of all individual classifiers;
[0082] Step 4.5 combines the voting results of each classifier in the previous Step 4.4 and selects the category ω with the most votes j as the final classification result of all data in cluster X.
[0083] Advantages of the present invention:
[0084] (1) The present invention proposes a semi-supervised classifier based on the nearest cluster. To ensure the accuracy of classification, the construction of the classifier relies on a quadratic clustering method based on K-Means. This classifier can handle both labeled data and unlabeled data simultaneously. It can not only effectively map the clustering results to specific industrial control protocols using limited labeled data, but also distinguish unknown industrial control protocol categories while identifying known industrial control protocol categories, thereby improving its flexibility and applicability in practical applications.
[0085] (2) The present invention is further improved on the basis of the proposed classifier and proposes a combined classification method based on a semi-supervised framework. This method is mainly implemented in two steps: First, the improved K-nearest neighbor density peak clustering algorithm is used to cluster industrial control traffic data; then, combined with the obtained clustering results and the prediction results of the classifier, a combined strategy is adopted to improve the accuracy of classification and ensure the reliability of the results. This method effectively integrates clustering and classification technologies, reducing the sensitivity of the classifier to errors and noises in the data while improving the recognition accuracy of industrial control traffic.
[0086] (3) Considering the scarcity of industrial control traffic data, the present invention aims to collect a large amount of industrial control traffic data at the data link layer based on traffic sessions in real time. By clustering and classifying with traffic sessions as the basic unit, this method can not only capture the subtle differences in industrial control network traffic in detail, but also effectively improve the accuracy of the algorithm. Description of the drawings
[0087] Figure 1 is a schematic diagram of the real-time cutting and collection process of industrial control traffic at the data link layer;
[0088] Figure 2 is a schematic diagram of a semi-supervised classifier based on the nearest cluster;
[0089] Figure 3 is a schematic diagram of the combined classification method;
[0090] Figure 4 Schematic diagram for performance comparison of different traffic classification technologies
[0091] Figure 5 Schematic diagram for comparison of different traffic classification technologies in identifying unknown industrial control protocol traffic Detailed implementation manners
[0092] In view of the problems faced by the research on industrial control network traffic classification, such as very limited traffic data, low accuracy of classification results, and large resource overhead, the present invention proposes an industrial control traffic classification method for the data link layer, which performs combined classification based on a semi-supervised framework. This method aims to bring together the respective advantages of different algorithms to improve the overall classification effect
[0093] Step 1: Complete the real-time cutting, collection, and preprocessing of industrial control traffic at the data link layer. The flow schematic diagram of the real-time cutting and collection of industrial control traffic at the data link layer is as Figure 1 shown
[0094] Loose start stage: The active time threshold Tact is initially set to a relatively high value of 120 seconds. Check the traffic sessions obtained after segmentation. If it is detected that the truncated traffic session is complete, it indicates that there is a large deviation between the current Tact value and the ideal truncation time, and the Tact value needs to be appropriately reduced. The formula is as follows
[0095] Tact = Tact / 2
[0096] Truncation avoidance stage: In this stage, it is considered that the Tact value has approached the actual ideal truncation time, and the Tact value should be reduced more cautiously to prevent incomplete traffic sessions caused by premature truncation. The update of Tact in this stage follows a linear rule of decrease, and the update formula is as follows
[0097] Tact = Tact - 5
[0098] Recovery stage: When it is found that the traffic session is incomplete during the inspection of the traffic session after truncation, the recovery stage is immediately entered. At this time, the truncation avoidance threshold Tthresh is set to 2 times the current Tact, and Tact is reset to 120 seconds, and then it returns to the loose start stage. This method continuously adjusts the active time threshold Tact by dividing stages, aiming to accurately locate the appropriate traffic session truncation time to optimize the cutting and collection process of traffic data
[0099] Data alignment: To determine the optimal value of the alignment length, that is, to minimize the sum of the interference information introduced by padding and the information lost due to truncation, this goal is evaluated by measuring the number of bits supplemented or lost due to padding and stage operations relative to the original data. The formula for the value range of L is defined as follows
[0100]
[0101] min(|x1|, …, |x n |) ≤ L ≤ max(|x1|, …, |x n |)
[0102] The principal component analysis (PCA) algorithm based on singular value decomposition (SVD) mainly includes data centering, singular value decomposition, principal component extraction and dimensionality reduction.
[0103] Data centering:
[0104] X c = X - μ
[0105] where the matrix X is an m×n matrix, where m is the number of samples and n is the number of features, and μ is the mean vector of each feature in the data matrix X.
[0106] Singular value decomposition: Perform singular value decomposition on the centered data matrix X c :
[0107] X c = U∑V T
[0108] where U is an m×m left singular matrix, ∑ is an m×n diagonal matrix containing singular values; V T is an n×n right singular matrix.
[0109] Principal component extraction and dimensionality reduction: The principal components can be extracted from the first k columns of V T where k is the number of principal components required, and the dimensionality-reduced data is represented as:
[0110] Z = X c V k
[0111] where, V k is the first k columns of the singular value matrix V T and Z is the dimensionality-reduced data matrix.
[0112] Step 2: Construct a semi-supervised classifier based on the nearest cluster, as shown in the schematic diagram Figure 2 as follows.
[0113] K-Means based clustering algorithm: Randomly select k points in the dataset as the initial centers, and assign each data point x i to the nearest cluster center; Calculate the mean of all points in each cluster as the new cluster center μ j , where S jis the set of all points in the j-th cluster under the current assignment; repeat the above process until the within-cluster sum of squares is minimized.
[0114]
[0115] According to the clustering results, apply a probability assignment mechanism to map the clusters created by K-Means to categories based on different industrial control protocols, and at the same time identify unknown industrial control protocols.
[0116] P(Y = y j |C i ) = n ij / n i
[0117] where the probability that the cluster C i is correctly mapped to the j-th category, where y j represents the traffic category, and n ij is the number of data points assigned to the cluster C i and whose category is j, and n i is the total number of labeled data points assigned to the cluster C i . According to the formula, the more the number of data points n i labeled as category j in the cluster C ij , the greater the probability that all data points in the cluster belong to category j.
[0118] Based on the posterior function, the data can be divided into three main categories, and different processing methods will be adopted for these categories later. For those clusters C i that do not contain any labeled data, they are defined as unknown categories. Set a threshold thre, and calculate the maximum value of the posterior function for each cluster. If the maximum value of the posterior function of a certain cluster C i exceeds the threshold thre, it means that all data points in the cluster C i can be classified into the same category, and the categories of all data points in the cluster are marked as y corresponding to the maximum value of the posterior function.
[0119]
[0120] For the data whose maximum value of the posterior function does not exceed the threshold thre, because their clustering attribution has high uncertainty, the present invention borrows the concept of fuzzy clustering to process them. By repeatedly executing the K-Means clustering step and performing re-clustering within the cluster, the classification accuracy is improved. After completing the secondary clustering, directly perform category division according to the new clustering results. At this time, it is no longer necessary to compare the maximum value of the posterior function with the threshold thre, but directly perform category assignment according to the corresponding criteria.
[0121] After completing the secondary clustering steps introduced in the previous section, recalculate the cluster centroids for all the generated clusters. The set of cluster centroids can be represented as M = {m1, m2, m3, …, m k′}, where k′ represents the number of clusters after secondary clustering. Depending on the actual situation, the same industrial control protocol may correspond to one or more clustering clusters. The industrial control protocol categories are represented by Y = {y1, y2, y3, …, y q}. For any category y i , as shown in the following formula, it is represented by the centroid set M i :
[0122] M i = {m j : C j ∈ y i}
[0123] The classification rule of the constructed classifier is defined as follows: for a given data point x, calculate the centroid m with the minimum distance to it, find the centroid set M i to which m belongs, and assign the label of the data point x to the category j represented by the centroid set M i . This classification formula follows the classification principle based on the nearest cluster.
[0124]
[0125] Step 3: Use the density peak clustering algorithm improved based on K-nearest neighbors (KNN-DPC) to cluster the industrial control traffic dataset;
[0126] Calculate the clustering matrix D of the dataset X using the Euclidean distance ij :
[0127]
[0128] For each point x i , calculate its k nearest neighbors in the set {x1, x2, …, x n} as follows: sort the Euclidean distances related to x i in the distance matrix in ascending order, and the distance at the k-th position determines the k-th nearest neighbor of x i . Denote the k nearest neighbors of x i as kNN(x i );
[0129] Sort({d i1 , d i2 , …, d iN})
[0130] kNN(x i ) = {j ∈ X|d(x i, x j ) ≤ d(x i , NN k (x i ))}
[0131] Calculate the local density ρ in the DPC algorithm i , draw on the core idea of the KNN algorithm, and propose a new calculation method for the local density ρ i
[0132]
[0133] Among them, the setting of k is based on a specific percentage of the total number of data points N, and its calculation formula can be expressed as:
[0134] k = p × N
[0135] Point x i The minimum distance δ i The calculation method can be defined according to the following two situations. If there is at least one point x j , whose density ρ j is greater than the density ρ of point x i , then δ i is defined as the minimum distance between point x i and all such points x i ; if the density of no point is greater than that of point x j , that is, x i is the highest density point, then δ i is defined as the maximum value of the distance between point x i and any other point in the dataset; i
[0136]
[0137] To optimize the selection of cluster centers, the present invention adopts a strategy of simultaneously considering the local density ρ i and the minimum distance value δ i , and selects points with both high local density values ρ i and far distances from other high-density points as cluster centers. For those points that are not selected as cluster centers, they are assigned to the cluster to which their nearest point with a higher density belongs, ensuring that each point belongs to a clear cluster.
[0138] Step 4: Combine the clustering results with the prediction results of the classifier to form a combined classification strategy, as shown in the schematic diagram Figure 3 as shown
[0139] Perform clustering on the dimension-reduced industrial control traffic dataset through Step 3; classify it unit by unit based on the individual traffic in the nearest cluster quadratic classifier through Step 2;
[0140] The combined classifier can be expressed by the following formula, where f(x) represents an individual classifier, F(x) represents the combined classifier, and θ is the combination function that merges the outputs of each individual classifier into the final classification decision
[0141] F(x) = θ x∈X (f(x))
[0142] Considering that different rules have different aggregation effects on classifiers and different sensitivities to noise data, it is finally decided to adopt the majority voting rule to summarize the prediction results of individual classifiers to determine the final classification decision
[0143]
[0144] According to the majority voting combination rule, all data x in cluster X i are set to the category ω j ;
[0145]
[0146] The experiments designed in the present invention apply four relatively advanced traffic classification technologies, namely C4.5, k-Means, and Naive Bayes (NB), to the field of industrial control networks, and conduct a comparative analysis of the performance of these methods with the combined classification method proposed in this paper.
[0147] To ensure the robustness of the experimental results and make full use of the data, the experiment uses ten-fold cross-validation as the method to evaluate the model performance. This method divides the dataset into ten subsets equally to ensure that each subset has the opportunity to be used as the test set once, while the other nine subsets are used as the training set. This process is repeated ten times. After ten iterations, the average value of the performance metrics obtained from these ten test cycles is calculated as an estimate of the overall performance of the model. Compared with the traditional method of dividing the dataset into a training set and a test set at one time, ten-fold cross-validation effectively reduces the dependence of the model evaluation results on a specific data partition by repeatedly using different data subsets as the training and test sets, thereby improving the accuracy and reliability of the model evaluation. In addition, it also allows the model to have the opportunity to use each data point in the dataset for training and testing, maximizing the utilization of the data. To comprehensively evaluate the experimental results, this study selects three key performance metrics: classification accuracy (CA), accuracy, and recall.
[0148] The experimental results are as Figure 4 and Figure 5 shown. It can be seen from Figure 4 that in all the investigated performance indicators, the method proposed in this paper is optimal compared with the other three algorithms. The method proposed in the present invention is superior to all other methods and achieves a higher accuracy rate. Figure 5 It is verified that the method proposed in the present invention has higher accuracy in dealing with unknown protocol traffic.
[0149] In summary, the two experiments designed in this section fully confirm that the combined classification method proposed in this paper can improve the classification effect in all performance indicators, and in dealing with the traffic data generated by unknown industrial control protocols, the combined classification method based on the semi-supervised framework proposed in the present invention shows significant advantages.
Claims
1. A data link layer-oriented industrial control traffic classification method, characterized in that: The steps include: Step 1: Cut, collect and preprocess the industrial control traffic at the data link layer in real time to obtain the data matrix Z after dimensionality reduction; Step 2: Build a semi-supervised classifier based on the nearest cluster, combine limited labeled data and a large amount of unlabeled data, and use secondary clustering technology to improve the reliability of clustering; Step 3: Propose a density peak clustering algorithm based on K-nearest neighbor improvement to cluster the industrial control flow data, calculate the local density and distance peak of each data point, and divide the industrial control flow data into different clusters; Step 4: Combine the classification results of the semi-supervised classifier based on the nearest cluster with the clustering results of step 3 to form a combined classification strategy and obtain the final classification result.
2. The data link layer-oriented industrial control traffic classification method according to claim 1 is characterized in that: The step 1 of real-time cutting and collection is specifically as follows: The real-time cutting and collection of industrial control traffic at the data link layer is divided into the loose start phase, the truncation start phase and the recovery phase; In the loose start phase, in order to obtain enough valid data while ensuring the rationality and effectiveness of traffic cutting, the active time threshold Tact is initially set to T1, 120≤T1≤150; the traffic session obtained after the industrial control traffic segmentation of the data link layer is checked, and if it is detected that the segmented and truncated traffic session is complete, it indicates that the current Tact value deviates from the ideal truncation time, and the Tact value is reduced; in this phase, the update of Tact follows the exponential law to decrease, and the formula is as follows: Tact=T / 2 When Tact decreases continuously and falls below the threshold Tthresh, it enters the truncation start phase; Tact updates in accordance with the linear law and decreases. The update formula is as follows: Tact=Tact-T2,(4≤T2≤8) After the industrial control traffic at the data link layer is truncated, if the traffic session is checked and found to be incomplete, the recovery phase is immediately entered; the truncation start threshold Tthresh is set to 2 times the current Tact, and Tact is reset to 120 seconds, and then returns to the loose start phase.
3. The data link layer-oriented industrial control traffic classification method according to claim 1 is characterized in that: The preprocessing in step 1 includes data alignment and principal component analysis based on singular value decomposition; The data alignment is evaluated by measuring the number of bits added or lost due to padding or truncation operations relative to the original data of the industrial control traffic at the data link layer; wherein the value range of the alignment length L is defined by the formula as follows; min(|x1|,…,|x n |)≤L≤max(|x1|,…,|x n |) Among them, x i represents the i-th binary industrial control traffic; The principal component analysis based on singular value decomposition includes data centering, singular value decomposition, principal component extraction and dimension reduction; The data set is specifically: X c =X-μ Among them, the data matrix X is an m×n matrix, where m is the number of samples, n is the number of features, and μ is the mean vector of each feature in the data matrix X; The singular value decomposition of the data matrix X after data centering c Perform singular value decomposition: X c =U∑V T Among them, U is an m×m left singular matrix, ∑ is an m×n diagonal matrix containing singular values; V T is an n×n right singular matrix; The principal component extraction and dimensionality reduction are specifically as follows: the principal component is extracted from the right singular matrix V T Extract V from the first f columns f , where f is the required number of principal components, the reduced data is represented as: Z=X c V f Among them, V f is the right singular value matrix V T The first f columns of , Z is the data matrix after dimensionality reduction.
4. The data link layer-oriented industrial control traffic classification method according to claim 3 is characterized in that: The optimal value of the alignment length L is the length value that minimizes the sum of the number of padding bits and the number of loss bits.
5. The data link layer-oriented industrial control traffic classification method according to claim 3 or 4, characterized in that: The semi-supervised classifier based on the nearest cluster is specifically: Step 2.1: Divide the reduced-dimensional data matrix into s clusters based on the K-Means clustering algorithm: randomly select s points in the data set as the initial centers, and for each data point x i , calculate its difference with all cluster centers μ j distance, find a cluster center that minimizes the distance, and place the data point x i Assigned to the set C corresponding to this cluster i For each cluster S j Calculate the coordinate mean of all points belonging to the cluster and use the mean as the new cluster center μ j ; Minimize the sum of squared distances from all points to the center of each cluster, so that the data points in the same cluster are as close as possible, thereby improving the clustering effect; Among them, x i represents a data point in the data set, μ j ∈{μ1,μ2,…μ s } represents all current cluster centers, d(x i , μ j ) represents x i To the center of mass μ j The distance of j Represents cluster S j The new center of mass of i Represents cluster C i The center of mass, ||x j -m i || represents point x j To the center of mass m i distance; Step 2.2: Based on the clustering results, a probability allocation mechanism is applied to map the clusters created by K-Means to categories based on different industrial control protocols, and unknown industrial control protocols are identified; P(Y=y j |C i )=n ij / n i Where P(Y=y j |C i ) represents cluster C i The probability of correctly mapping to the jth category, y j Indicates the traffic category, n ij is assigned to cluster C i The number of data points in the class j, n i is assigned to cluster C i The total number of labeled data points; cluster C i The data point n that is labeled as category j in ij The larger the number, the greater the probability that all data points in the cluster belong to category j; Step 2.3: Based on the posterior function, the preprocessed data is divided into three main categories: unknown class, known class, and fuzzy class; for cluster C that does not contain any labeled data i , defining it as an unknown class; clusters containing labeled data are judged to belong to known classes or fuzzy classes based on the posterior probability; Step 2.4, set the threshold thre and calculate the maximum value of the posterior function of each cluster; when a cluster C i If the maximum value of the posterior function exceeds the threshold thre, cluster C i All data points in are classified into the same category, and the categories of all data points in the cluster are marked as the traffic category y corresponding to the maximum value of the posterior function; Step 2.5: For data whose maximum value of the posterior function does not exceed the threshold value thre, the fuzzy clustering concept is used to process it; by repeating the K-Means clustering step, the cluster is re-clustered; after completing the secondary clustering, the category division is directly performed according to the new clustering results, and the category allocation is directly performed according to the corresponding criteria; Step 2.6: After completing the secondary clustering, recalculate the cluster centers of all generated clusters; the set of cluster centers is represented by M = {m1, m2, m3, ..., m k′ }, k′ represents the number of clusters after secondary clustering; the same industrial control protocol will correspond to one or more clusters; the industrial control protocol category is represented by Y = {y1, y2, y3, …, y q } to indicate that for any category y i , through the centroid set M of sample points belonging to this category i To express; M i ={m j :C j ∈y i } Step 2.7: The classification rule of the semi-supervised classifier based on the nearest cluster is defined as: for a given data point x, calculate the centroid m with the smallest distance to it, and find the centroid set M to which m belongs. i , the label of the data point x is set as the centroid set M i The category j represented; 6. The data link layer-oriented industrial control traffic classification method according to claim 5 is characterized in that: The density peak clustering method based on K nearest neighbor improvement is specifically as follows: Step 3.1: Use Euclidean distance to calculate the distance matrix D of the dataset X ij ; Among them, d(x i ,x k ) represents x i and x k The distance between ij and x kj Represents the data point x i and x k The value of the jth feature; Step 3.2: For each point x i , in the set {x1, x2, …, x n } calculate its k nearest neighbors; The distance matrix D ij Zhong and x i The relevant Euclidean distances are arranged in ascending order, and the distance at the kth position is taken as x i The kth nearest neighbor of i ) represents x i The k nearest neighbors of ; Sort({d i1 ,d i2 ,…,d iN }) kNN(x i )={j∈X|d(x i ,x j )≤d(x i ,NN k (x i ))} Among them, d(x i , x j ) represents x i and x j The distance between them, NN(x i ) represents x i The kth nearest neighbor of , and kNN(x i ) is a set containing all the i The point within the kth nearest neighbor range; Step 3.3: Calculate the local density ρ in the density peak clustering algorithm i , the local density ρ i The calculation method is as follows: The setting of k is based on a specific percentage of the total number of data points N, and its calculation formula is expressed as: k=p×N Where N represents the total number of data points in the data set, and p represents a percentage value between 0 and 1; Step 3.4, click x i The minimum distance δ i The calculation method is defined according to the following two cases: when there is at least one point x j , its density ρ j Greater than point x i The density ρ i , then δ i is defined as the point x i With all points x j The minimum distance between points; when no point has a density greater than point x i , then point x i is the point with the highest density, δ i is defined as the point x i The maximum distance to any other point in the data set; Step 3.5: Consider the local density ρ i and the minimum distance value δ i The strategy is to select a high local density value ρ i , and the points that are far away from other high-density points are used as cluster centers; for those points that are not selected as cluster centers, they are assigned to the cluster to which their nearest points with higher density belong, ensuring that each point belongs to a clear cluster.
7. The data link layer-oriented industrial control traffic classification method according to claim 6 is characterized in that: The step 4 is specifically as follows: Step 4.1: Through step 2, each piece of industrial control flow data is input into a semi-supervised classifier based on the nearest cluster for classification, and each piece of flow data is divided into more refined categories; Step 4.2: Perform preliminary clustering of the industrial control traffic data set using the K-nearest neighbor-based improved density peak clustering algorithm (KNN-DPC) on the industrial control traffic data after dimensionality reduction in step 3; Step 4.3: Integrate the output results of each individual classifier through the combination function θ to form the final output result F(x) of the combined classifier; F(x)=θ x∈X (f(x)) Among them, F(x) represents the output result of the combined classifier, which represents the final classification decision for sample x, and f(x) represents the output result of a single classifier, that is, the prediction result of each individual classifier for sample x; Step 4.4: Use the majority voting rule to aggregate the prediction results of individual classifiers to determine the final classification decision; Among them, v ij Used to record the i-th single classifier for category ω j Voting status, y xi Represents the predicted category of sample x by the i-th classifier; Step 4.5: Use the majority voting rule to select the category with the most votes from all classifiers as the final classification result of the sample; in, Represents category ω j The total number of votes is the sum of the votes of all individual classifiers; Step 4.5 combines the voting results of each classifier in the previous step 4.4 and selects the category with the most votes ω j As the final classification result of all data in cluster X.