An Internet-connected industrial control asset identification method based on an improved decision tree

By improving the decision tree model and Adaboost algorithm, combined with the characteristics of industrial control equipment, the efficiency and accuracy of industrial control equipment asset identification in the industrial control field are solved, and efficient and stable industrial control equipment identification is achieved.

CN116318936BActive Publication Date: 2025-07-04NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310212754.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2025-07-04
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

The prior art lacks effective methods in the field of industrial control to efficiently and granularly identify networked industrial control equipment assets. Traditional detection methods are prone to affect the functional safety of the equipment and have noise interference, making it difficult to meet the dual needs of functional safety and information security.

Method used

Using a method based on improved decision tree, combined with the characteristics of industrial control equipment, through the networked industrial control asset detection model, fingerprint feature combination and feature weight correction, a decision tree model is constructed and the Adaboost algorithm is used to improve the recognition accuracy and generalization ability.

Benefits of technology

It realizes efficient and stable identification of the operating system and model of industrial control equipment, reduces sample dependence, reduces noise interference, and improves identification accuracy and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116318936B_ABST
    Figure CN116318936B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of network security and proposes a method for identifying networked industrial control assets based on an improved decision tree. This method uses a networked industrial control asset detection model to conduct data detection to form an industrial control device traffic data set; proposes a fingerprint feature combination, which is combined with the industrial control device traffic data set for preprocessing to form a data set; divides the data set into a networked industrial control device operating system asset data set and a networked industrial control device model asset data set according to data labels; calculates feature weights; constructs a decision tree model corrected based on feature weights, and establishes a decision tree method based on Adaboost according to the constructed decision tree model to improve accuracy and reduce coupling. The present invention improves the accuracy and generalization ability of the identification method, and has higher accuracy, precision and better coverage rate whether for the operating system or the device model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security, and in particular, to a method for identifying networked industrial control assets based on an improved decision tree. Background Art

[0002] As a connection point for connecting IT (Information Technology) and OT (Operation Technology), the industrial Internet has currently been widely applied to industrial control monitoring systems in various industries such as power operation, oil and gas transportation, chemical synthesis, and transportation. While networked industrial control devices bring convenience to the monitoring of industrial sites, they also introduce possible risks existing in network security into the industrial production environment. Under the guidance of the "dual security integration" of functional safety and information security, the implementation of industrial control network security must ensure the normal operation of the functions of industrial site equipment. Therefore, research on active defense and situation awareness technologies for networked industrial control devices is very important.

[0003] Comprehensively and meticulously mastering the asset information of industrial control devices can facilitate the software and hardware upgrades of the devices, update vulnerability information, and help security managers reduce the pressure of security alerts and formulate preventive strategies in advance for possible network attacks targeting specific devices or systems, enabling the security defense of the industrial Internet to shift from passive to active. Although there are currently some basic researches on network asset identification and the research in the field of industrial network security is also increasing day by day, there is still a lack of effective means for the problem of identifying networked industrial control device assets, and the fingerprint features of industrial control devices are still a blank area. There is still no effective method to efficiently solve the problem of identifying networked industrial control device assets using machine learning methods.

[0004] The paper "Shen Yu. Research and Application of Network Asset Automatic Identification Method [D]. Shanxi University, 2021. DOI: 10.27284 / d.cnki.gsxiu.2021.001759." proposes a method for identifying operating systems based on CNN, using the p0f fingerprint database and traffic data collected by network asset detection as the data set to train the CNN model to improve the accuracy of identifying operating systems. This method still only targets operating systems, and the characteristic attributes in the fingerprint database still rely on the fingerprint database of the p0f tool without being adjusted in combination with the characteristics of actual devices.

[0005] The paper "Yang Yan, Liu Jianhua, Tian Dongping. Identification of Information Assets Based on Decision Tree Algorithm [J]. Modern Electronics Technique, 2010, 33(23): 77-79+84. DOI: 10.16652 / j.issn.1004-373x.2010.23.048." proposed an identification method based on the C4.5 decision tree algorithm. Based on the TCP / IP protocol family, at the network layer, different operating systems have different settings for options in network communication. The combination status of options is divided by the C4.5 decision tree to complete the classification of operating systems. This research verified the effectiveness of the decision tree algorithm in the field of network asset identification. However, it did not adjust the algorithm details according to specific problems and data characteristics. Moreover, the selection of the dataset only targeted the network layer communication protocol of traffic data, and the data performance was not comprehensive enough.

[0006] Currently, the research on asset identification mainly stays at the personal PC side in the network, and mainly uses some basic scanning tools such as Nmap to scan and count information such as operating systems and open ports in network assets, or relies on existing documents for typical device identification based on document retrieval. There is not much research on the identification of networked industrial control devices in special scenarios in the industrial control field, and there is a lack of effective methods in related fields. Due to the existence of a large amount of asset information that cannot be detected and obtained in the industrial environment, using traditional detection methods is likely to cause the function of industrial control devices to stagnate due to a large amount of traffic injection, and there are honeypot devices interfering with asset information identification. Therefore, how to complete high-quality, efficient, and fine-grained detection of industrial control field devices and use the unique features of industrial control devices combined with effective machine learning means to complete efficient, low-frequency, and noise-resistant identification of networked industrial control assets while ensuring functional safety and information security is a daunting challenge. Summary of the Invention

[0007] The present invention combines the characteristics of industrial control devices, analyzes the protocols applied to the industrial control field, and proposes a method for identifying networked industrial control assets based on an improved decision tree for the industrial Internet. This method is used to train the operating system and industrial control device model datasets in the industrial control environment respectively. This method aims at the problem of how to analyze and classify traffic information when device information cannot be directly obtained through detection methods. In solving the problem of identifying networked industrial control assets, this method has the characteristics of stability, high efficiency, reduced sample dependence, being able to adapt to the characteristics of traffic data, and being able to solve classification problems.

[0008] The technical solution of the present invention is as follows: A method for identifying networked industrial control assets based on an improved decision tree, including the following steps;

[0009] Step 1: Use a networked industrial control asset detection model to detect data; the detected information forms an industrial control device traffic dataset;

[0010] The Internet-connected industrial control asset detection model is a comprehensive industrial Internet detection model. This model combines active detection, passive detection, detection based on search engines, industrial control communication protocols, and the detection method of industrial control honeypot identification to detect the operating system and device model asset information of industrial control devices connected to the Internet. The specific workflow is as follows;

[0011] 1.1. Input the IP address of the target host;

[0012] 1.2. Send an ICMP protocol data packet. If the target host responds with an Echo reply message, go to step 1.3; if no response from the target host is received within the specified time, this round of detection ends, and step 1.1 is restarted to input a new target host IP address;

[0013] 1.3. Send a SYN data frame to the target port of the target host. When the target host replies with a SYN ACK data frame, proceed to step 1.4; if the target host does not respond, this round of detection ends, and step 1.1 is restarted to input a new target host IP address;

[0014] 1.4. Select an industrial control protocol data message according to the open port situation of the target host and send it, and monitor the traffic of the target host; when the reply message from the target host conforms to the industrial control protocol standard, proceed to step 1.5; when the target host does not respond or responds with an error, this round of detection ends, and step 1.1 is restarted to input a new target host IP address;

[0015] 1.5. Check the honeypot data characteristics of the traffic data of the target host, including the target host service provider, the situation of multiple open ports, and the length of the response time. When the target host does not have the above honeypot data characteristics, proceed to step 1.6; when the target host has any of the above honeypot data characteristics, proceed to step 1.9;

[0016] 1.6. Use the shodan search engine interface to check the information of the target host. There are three situations in total; when shodan marks the target host as a suspected honeypot, proceed to step 1.7; when shodan marks the target host as a confirmed honeypot, proceed to step 1.9; when it does not meet the above two situations, the target host is considered an industrial control device, and proceed to step 1.10;

[0017] 1.7. After waiting for 30 seconds, resend the industrial control protocol connection request and update the industrial control protocol function command code;

[0018] 1.8. Check the communication behavior of the target host. When there is a connection refusal or the new industrial control protocol function command code is not supported, it is considered that the target host has the behavior characteristics of a honeypot and proceed to step 1.9, otherwise proceed to step 1.10;

[0019] 1.9 Mark the target host as a honeypot device, discard the traffic data of the target host, and end the detection;

[0020] 1.10 Mark the target host as an industrial control device, save the traffic data of the target host, and store the relevant detection information in the industrial control device traffic dataset.

[0021] Step 2: Propose a fingerprint feature combination, combine the industrial control device traffic dataset with this fingerprint feature combination, perform data preprocessing through normalization or one-hot encoding to form a dataset; obtain data labels according to the detection information, and divide the dataset into an operating system asset dataset of networked industrial control devices and a model asset dataset of networked industrial control devices according to the data labels;

[0022] The fingerprint feature combination includes the conventional fingerprint features in TCP / IP protocol data packets and the fingerprint features selected according to the characteristics of industrial control devices themselves; among the mean, variance, range, maximum value, and minimum value of the variable data in the industrial control device traffic dataset, the fields with regular changes are subjected to data reduction and data aggregation to describe the periodic changes of the data stream characteristics; therefore, the fingerprint features selected according to the characteristics of industrial control devices themselves include the dimension num, mean av, variance var, maximum value max, minimum value min, extreme value mr, port port, industrial control protocol service, whether to use industrial Ethernet ifet, longitude longitude, and latitude latitude; the conventional fingerprint features include whether to use the TCP protocol ty at the transport layer, whether to use the UDP protocol uy at the transport layer, sliding window length tws, whether to allow fragmentation df, target host response time ptime, indication of congestion handling ecn, differentiated service area dsf, whether it is the last fragmentation bit mf, device special identifier lg, distinction between unicast or multicast ig, packet length pl, TCP / UDP header length tul, and time to live ttl;

[0023] Combine the industrial control device traffic dataset with the fingerprint feature combination to obtain the combined dataset;

[0024] 2.1 For the numerical fingerprint features in the combined dataset, calculate their variance, mean, and range;

[0025] Set the fingerprint feature A in the combined dataset i There are n values in total, where a ij is the jth value of the fingerprint feature A i ; the mean av(A i ) in the fingerprint feature A i is defined as follows:

[0026]

[0027] Fingerprint feature A iThe variance of var(A i ) is defined as follows:

[0028]

[0029] The fingerprint feature A i The range mr(A i ) is defined as follows:

[0030] mr(A i ) = max(A i ) - min(A i )

[0031] 2.2. Normalize the numerical fingerprint features in the combined dataset that do not conform to the standard normal distribution, so that the overall mathematical distribution of this type of numerical fingerprint features conforms to the standard normal distribution, reduce the magnitude of these data, and avoid the change of large-value features masking small-value features; Use Z-score normalization, and the normalization conversion process is shown in the following formula, where A is the fingerprint feature set and a i is the specific fingerprint feature value:

[0032]

[0033] 2.3. For the discrete categorical features in the combined dataset that cannot be infinitely subdivided and the number of values is countable, perform one-hot encoding to make the fingerprint feature values of this type ordered and continuous;

[0034] 2.4. According to the detection information collected in step 1.10, use regular expressions to extract the valid information and use it as the data label. According to the types of data labels, divide the dataset processed in steps 2.1 - 2.3 into the networked industrial control device operating system asset dataset and the networked industrial control device model asset dataset.

[0035] Step 3. Since the dataset comes from the real industrial control device traffic, there are differences between the data sample features; considering that the different features of the industrial control traffic dataset have different importance for the asset identification and classification problem, and at the same time to solve the differences between the sample data fingerprint features in the networked industrial control device operating system asset dataset and the networked industrial control device model asset dataset, calculate the feature weights to correct the Gini coefficient;

[0036] Perform the following operations on the networked industrial control device operating system asset dataset and the networked industrial control device model asset dataset respectively, specifically;

[0037] 3.1. Randomly select a sample R l from the networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset, requiring that l is less than or equal to the input random sampling number O, for each sample R lPerform operations from step 3.2 to step 3.3;

[0038] 3.2. Calculate the sample R l and the Euclidean distance between R l and each sample of the same category as R. Then find the set H of the k nearest neighboring samples of the same category. Since the collected data set exists in the form of a sample set, the same category here actually refers to the samples that belong to the same sample set as the sample R l ;

[0039] The Euclidean distance is calculated using the following formula,, where x α represents the coordinates of the sample point P, and y α represents the coordinates of the sample point B;

[0040]

[0041] P and B are respectively sample points in the data set D; c is the dimension of the data points;

[0042] 3.3. Calculate the Euclidean distance between the sample R l and each sample of a different category from R l and also find the set M of the k nearest neighboring samples of different categories;

[0043] 3.4. Traverse the set H of samples of the same category as the sample R l and perform the operation of step 3.5 for each H p in the set H of samples of the same category;

[0044] 3.5. When the fingerprint feature A i has different values on the sample R l and the sample H p , then diff_rh is incremented by 1. diff_rh represents the intermediate value when calculating the feature weight of samples of the same category;

[0045] 3.6. Calculate the feature weight value w_same of the fingerprint feature A l in the set H of samples of the same category as the sample R i according to the following formula, where m is the number of iterations:

[0046]

[0047] 3.7. Traverse the set CM of categories in the set M of samples of different categories from the sample R l and perform the operations of steps 3.8 to 3.10 for each category C q in CM;

[0048] 3.8. Traverse the set M of samples of different categories from the sample R l with the category C qThe sample set M(C q ), for each sample M in M(C q ), perform operation 3.9; x Perform operation 3.9;

[0049] 3.9. When the fingerprint feature A i has different values on the sample R l and the sample M x , then diff_rm is incremented by 1. diff_rm represents an intermediate value when calculating the feature weights of heterogeneous samples;

[0050] 3.10. Calculate the weight value w_diff of the fingerprint feature A in the heterogeneous sample set M of the sample R l according to the following formula, where s is the number of samples in M(C i ), class(R) represents the class of the sample R, and P represents the probability: q w_diff += P(C

[0051] w_diff += P(C q ) * diff_rm / (1 - P(class(R)) * s * k)

[0052] 3.11. Calculate the feature weight W(A i ) of each fingerprint feature A in the fingerprint feature set A: i W(A

[0053] W(A i ) = W(A i ) - w_samw + w_diff.

[0054] Step 4. Construct a decision tree model based on the corrected feature weights. Use the product of the feature weights calculated in Step 3 and the Gini coefficient as the node splitting criterion to establish the decision tree model; consider the importance of the fingerprint feature and the accuracy of the model after classifying the fingerprint feature when establishing the decision tree model;

[0055] 4.1. Initialize and create a root node that contains the entire networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset;

[0056] 4.2. Multiply each feature weight W(A i ) in the feature weight set W(A) calculated in Step 3 by the Gini coefficient Gini(D) to obtain the corrected Gini coefficient, which is used as the basis for selecting the feature for node splitting;

[0057] The Gini coefficient Gini(D) of the fingerprint feature A for the networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset is calculated by the following formula;

[0058] Q1 = {x ∈ D | A(x) = a}

[0059] Q2 = D - Q1

[0060]

[0061]

[0062] Where Q1 is the set of samples in the operating system asset dataset of networked industrial control devices or the model asset dataset of networked industrial control devices where the fingerprint feature A takes the value a; Q2 is the set of samples in the operating system asset dataset of networked industrial control devices or the model asset dataset of networked industrial control devices where the fingerprint feature A does not take the value a; assume the operating system asset dataset of networked industrial control devices or the model asset dataset of networked industrial control devices is D i The class set E on the fingerprint feature A is divided into {E1, E2, ……, E n};

[0063] 4.3. Determine to select the fingerprint feature A i as the node division criterion; the operating system asset dataset of networked industrial control devices or the model asset dataset of networked industrial control devices is divided into several subsets D i according to different values of the fingerprint feature A i , and delete the fingerprint feature A in the fingerprint feature set A i ;

[0064] 4.4. For each subset D i in the subset D ij , perform operations from step 4.5 to step 4.7;

[0065] 4.5. When all instances in D ij belong to the same class, create a leaf node and label the node with the class;

[0066] 4.6. When D ij is empty, return to step 4.4;

[0067] 4.7. When the number of nodes in D ij is less than the number of neighboring samples k + 1 and the number of subset nodes is less than the random sampling number f + 1, create a leaf node and label the node with the class having the largest number in D ij ;

[0068] 4.8. When D ij does not meet all the above conditions, go to step 4.2;

[0069] 4.9. Output the constructed decision tree model based on feature weight correction.

[0070] Step 5. Combine the decision tree model based on feature weight correction constructed in Step 4 to establish a decision tree method based on Adaboost, improving accuracy and reducing coupling;

[0071] 5.1. Define a set of fingerprint feature vectors U = {u1, u2, u3,..., u t} for the operating system asset dataset or model asset dataset of networked industrial control devices. The corresponding set of classification labels for U is defined as V = {v1, v2, v3,..., v t}. Specify that the weight of each sample u in the feature vector set is the same, which is i.e., the initial feature weight Input the number of training loops N. Each loop executes Steps 5.2 to 5.9;

[0072] 5.2. The weight coefficient α h of the classifier model in the h-th iteration is initially set to 0. α h represents the weight of the weak classifier in the h-th iteration in the final strong classifier. Extract the training set T h from the set of fingerprint feature vectors U, which contains r samples;

[0073] 5.3. Use the decision tree model based on feature weight correction proposed in Step 4 to train the weak classifier model C h in the h-th iteration, as shown by the following formula:

[0074] C h = s_Dtree(T h )

[0075] 5.4. For each sample u h in the training set T g , perform the operations in Steps 5.5 to 5.7;

[0076] 5.5. When the classification result C g of the weak classifier model in the h-th iteration for the sample u h (u g ) is different from the original class label v g of the sample u h , it proves that the classification by type fails. The error rate e h of the classifier model in the h-th iteration is calculated as follows, where w h (g) represents the weight of the g-th sample in the training set T h :

[0077] e h += w h (g)

[0078] 5.6. To ensure that each weak classifier has better classification performance than a random classifier, it is required that the error rate e h is less than 0.5. When e h is greater than or equal to 0.5, end this training, go to step 5.2, and start a new round of loop;

[0079] 5.7. When the error rate e h is less than 0.5, then calculate α h . The formula is as follows:

[0080]

[0081] 5.8. When the classification result C g of the h-th round of iterative weak classifier model for the sample u h (u g ) is different from the original class label v g of the sample u g , to ensure that the new round of weights calculated by the weight calculation formula conform to the probability distribution, calculate the normalization factor Z h . The formula is as follows:

[0082]

[0083] The weight value w g of the sample u h+1 used in the new round of iteration is calculated as follows:

[0084]

[0085] 5.9. When the classification result C g of the h-th round of iterative weak classifier model for the sample u h (u g ) is the same as the original class label v g of the sample u g , it proves that the classification is successful. The formula for Z h is as follows:

[0086]

[0087] The formula for w h+1 (g) is as follows:

[0088]

[0089] 5.10. After all training loops are completed, return the strong classifier f(x). The formula is as follows:

[0090] f(x) = ∑α h C h ..

[0091] The beneficial effects of the present invention are as follows:

[0092] (1) The proposed networked industrial control asset detection model abstracts and defines the finite states of the model. In the host liveness detection stage, the ICMP protocol is selected for liveness detection, and in the port liveness detection stage, the half-connection type of TCP SYN is selected for liveness detection, reducing the number of packets sent and improving the detection efficiency. The detection model uses the shodan interface and the behavioral and data characteristics of honeypots to help reduce noise, avoiding the impact of excessive injection of real industrial control traffic on device real-time performance, increasing the authenticity of detection data, and saving available traffic data for analysis and processing;

[0093] (2) Combining the detection data of the networked industrial control asset detection model with the characteristics of industrial control devices, a new fingerprint combination of industrial control network traffic data is extracted based on the fingerprints of traditional network traffic data; the traffic data is processed according to the proposed fingerprint combination of industrial control network traffic data to form an industrial control asset fingerprint dataset. Using the dataset as input, the effectiveness of the fingerprint combination of industrial control network traffic data in the field of industrial control asset identification is verified using a decision tree model;

[0094] (3) A decision tree model based on feature weight correction is proposed in combination with the characteristics of industrial control traffic data, and a decision tree method based on Adaboost is established. This method calculates the feature weights and corrects the drawback of the Gini coefficient that only cares about the effect of randomly selecting a sample that is misclassified after classifying the sample for a feature, while ignoring the importance of the sample for classification;

[0095] (4) The decision tree model based on feature correction is strengthened based on the Adaboost algorithm to improve the accuracy and generalization ability of the recognition method, and the effectiveness is verified through comparative experiments. Description of the Drawings

[0096] Figure 1 is a flowchart of a networked industrial control asset recognition method based on an improved decision tree;

[0097] Figure 2 is a flowchart of a networked industrial control asset detection model;

[0098] Figure 3 is a statistical chart of the number of operating systems of industrial control devices;

[0099] Figure 4 is a comparative experimental chart of the detection effect of networked industrial control devices;

[0100] Figure 5 is a comparative chart of the recognition effect of the operating systems of networked industrial control devices;

[0101] Figure 6 is a comparative chart of the recognition effect of the models of networked industrial control devices. Detailed Implementation Modes

[0102] The specific implementation manners of the present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings.

[0103] The present invention proposes a method for identifying networked industrial control assets based on an improved decision tree. Figure 1 The flowchart of a method for identifying networked industrial control assets based on an improved decision tree proposed by the present invention is shown. It mainly includes a networked industrial control asset detection model and a decision tree method based on Adaboost. The steps of the whole method are as follows:

[0104] The specific implementation steps of the whole method for identifying networked industrial control assets based on an improved decision tree are as follows:

[0105] Step 1: Use the networked industrial control asset detection model for data detection, and the specific implementation is as follows:

[0106] 1.1. Input the IP address of the target host. First, it is necessary to ensure that the target host is in a communicable state, that is, the target host is connectable, communicable, and has feedback on the Internet. Entering the host alive state can ensure that in the subsequent communication detection process, the feedback of the target host can be received;

[0107] 1.2. Send ICMP protocol data packets. If the target host responds with an Echo reply message, go to 1.3. If no response is received within the specified time, the current round of detection ends, and step 1.1 is restarted to input a new target host IP address;

[0108] 1.3. Send a SYN data frame to the target port of the target host. If the target host replies with a SYN ACK data frame, go to 1.4. If there is no response, the current round of detection ends, and step 1.1 is restarted to input a new target host IP address;

[0109] 1.4. Select and send industrial control protocol data packets according to the port opening situation, and monitor the target host traffic. If the packet replied by the target host conforms to the industrial control protocol standard, go to 1.5. If there is no response or the response is incorrect, the current round of detection ends, and step 1.1 is restarted to input a new target host IP address;

[0110] 1.5. Check the honeypot data characteristics of the target host traffic data, including the target host service provider, multi-port opening situation, and response time length. If there are no above honeypot data characteristics, go to 1.6. If the above honeypot data characteristics are available, go to 1.9;

[0111] 1.6 Use the Shodan search engine interface to check the IP information of the target host. There are three labeling statuses of Shodan for honeypots: confirm that the target host is a honeypot device, the target host is suspected of being a honeypot, and the target host is a non-honeypot device. Based on this, there will be different state transitions. If Shodan marks the host IP as a suspected honeypot, go to 1.7; if it is marked as a confirmed honeypot, go to 1.9. When neither of the above two cases is met, the target host is considered an industrial control device, and proceed to step 1.10;

[0112] 1.7 After waiting for 30 seconds, resend the industrial control protocol connection request and update the industrial control protocol function command code;

[0113] 1.8 Check the communication behavior of the target host. If it has a behavior characteristic of refusing to connect or not supporting new function command codes, it is considered that the target host has the behavior characteristics of a honeypot, go to 1.9, otherwise go to 1.10;

[0114] 1.9 Mark the target host as a honeypot device, discard the traffic data of the target host, and end the detection;

[0115] 1.10 Mark the target host as an industrial control device, save the traffic data of the target host, and store the relevant detection information into the industrial control device traffic data set.

[0116] Step 2: Establish a data set and process each fingerprint feature as follows: The fingerprint feature combinations include the conventional fingerprint features in Table 1 and the fingerprint features in Table 2 selected according to the characteristics of industrial control devices themselves.

[0117] Table 1 TCP / IP Conventional Fingerprint Features

[0118]

[0119] Table 2 Fingerprint Features Selected According to the Characteristics of Industrial Control Devices Themselves

[0120]

[0121] 2.1 Calculate the variance, mean, and range of numerical features such as ttl, tws, pl, and tul;

[0122] 2.2 Assume that a fingerprint feature of a data is A i There are n values in total, where a ij is the jth value of feature A i . The mean av(A i ) of fingerprint feature A i is defined as follows:

[0123]

[0124] 2.3 Fingerprint feature A iThe variance var(Ai) is defined as follows:

[0125]

[0126] 2.4 Fingerprint feature A i The range mr(Ai) is defined as follows:

[0127] mr(A i ) = max(A i ) - min(A i )

[0128] 2.5 Normalize numerical features such as ttl, tws, and pl so that these data reduce the magnitude and avoid large-value feature changes from masking small-value features. Use Z-score normalization, and the normalization conversion process is shown in the following formula, where A is the fingerprint feature set and a i is the specific feature value:

[0129]

[0130] 2.6 Perform one-hot encoding on discrete categorical features such as service and port;

[0131] 2.7 The banner field in the collected probe data contains probe information. Valid information can be extracted using regular expressions and used as data labels.

[0132] Table 3 Networked industrial control device models

[0133]

[0134] The data labels divide the data into a networked industrial control device operating system asset dataset and a networked industrial control device model asset dataset, as shown in Figure 3 and Table 3 respectively.

[0135] Step three: Calculate the feature weights. The calculation process is as follows:

[0136] 3.1 Randomly select a sample R l from the networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset, requiring l to be less than or equal to the input random sampling number O. Perform operations from step 3.2 to step 3.3 on each sample R l ;

[0137] 3.2 Calculate the Euclidean distance between sample R l and each sample of the same category as R l , and find the set H of k nearest neighboring samples of the same category. The Euclidean distance is calculated using the following formula, where x α represents the coordinates of sample point P, and yα Denote the coordinates of sample point B;

[0138]

[0139] P and B are respectively sample points in the data set D; c is the dimension of the data points;

[0140] 3.3. Calculate sample R l and the Euclidean distances of each sample with a different category from R l Also, find the set M of the k nearest neighboring samples of different categories;

[0141] 3.4. Traverse the set H of the same-category samples of sample R l For each H in the set H of the same-category samples of sample R p Perform the operation in step 3.5;

[0142] 3.5. When the fingerprint feature A i has different values on sample R l and sample H p increase diff_rh by 1, where diff_rh represents the intermediate value when calculating the feature weight of the same-category samples;

[0143] 3.6. Calculate the feature weight value w_same of the fingerprint feature A l in the set H of the same-category samples with the same category as sample R i according to the following formula, where m is the number of iterations:

[0144]

[0145] 3.7. Traverse the set CM of categories in the set M of different-category samples of sample R l For each category C in CM q Perform the operations from step 3.8 to step 3.10;

[0146] 3.8. Traverse the sample set M(C l ) with category C in the set M of different-category samples of sample R q For each sample M in M(C q ) q perform the operation in step 3.9; x

[0147] 3.9. When the fingerprint feature A i has different values on sample R l and sample M x increase diff_rm by 1, where diff_rm represents the intermediate value when calculating the feature weight of different-category samples;

[0148] 3.10. Calculate the following formula for sample R​l The fingerprint feature A in the heterogeneous sample set M i The weight value w_diff, where s is the number of samples in M(C q ), class(R) represents the class of sample R, and P represents probability:

[0149] w_diff += P(C q ) * diff_rm / (1 - P(class(R)) * s * k)

[0150] 3.11. Calculate the feature weight W(A i ) of each fingerprint feature A in the fingerprint feature set A i ):

[0151] W(A i ) = W(A i ) - w_same + w_diff.

[0152] Step Four: Construct a decision tree based on feature weight correction, and use the product of the feature weight calculated in Step Three and the Gini coefficient as the basis for selecting features for node splitting. The specific steps are as follows:

[0153] 4.1. Initialize and create a root node containing the entire networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset;

[0154] 4.2. Multiply each feature weight W(A i ) in the feature weight set W(A) calculated in Step Three by the Gini coefficient Gini(D) to obtain the corrected Gini coefficient, which is used as the basis for selecting features for node splitting;

[0155] The Gini coefficient Gini(D) of the fingerprint feature A for the dataset D is calculated by the following formula;

[0156] Q1 = {x ∈ D | A(x) = a}

[0157] Q2 = D - Q1

[0158]

[0159]

[0160] where Q1 is the set of samples in the dataset D where the fingerprint feature A takes the value a; Q2 is the set of samples in the dataset D where the fingerprint feature A does not take the value a; it is assumed that the class set E of the dataset D i on the fingerprint feature A is divided into {E1, E2, ……, E n};

[0161] 4.3. Determine to select the fingerprint feature A iis the node division criterion; the networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset is divided into several subsets D according to different values of fingerprint feature A i ; delete fingerprint feature A from the fingerprint feature set A i ; i ;

[0162] 4.4. For each subset D i in the subset D ij , perform operations from step 4.5 to 4.7;

[0163] 4.5. When all instances in D ij belong to the same class, create a leaf node and label the node with the class;

[0164] 4.6. When D ij is empty, return to step 4.4;

[0165] 4.7. When the number of nodes in D ij is less than the number of neighboring samples k + 1 and the number of subset nodes is less than the random sampling number f + 1, create a leaf node and label the node with the class having the largest number in D ij ;

[0166] 4.8. When D ij does not meet all the above conditions, go to step 4.2;

[0167] 4.9. Output the constructed decision tree.

[0168] Step Five: Combine the decision tree constructed in Step Four to establish a decision tree method based on Adaboost to improve accuracy and reduce coupling. The steps of the algorithm are as follows:

[0169] 5.1. Define a set of fingerprint feature vectors U = {u1, u2, u3,..., u t} for the networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset. The corresponding classification label set of U is defined as V = {v1, v2, v3,..., v t}. Specify that the weight of each sample u in the feature vector set is the same, which is , that is, the initial feature weight . Input the number of training loop times N. Each loop executes steps 5.2 to 5.9;

[0170] 5.2. The weight coefficient α h of the classifier model in the h-th iteration is initially set to 0. α h represents the weight of the weak classifier in the h-th iteration in the final strong classifier. Extract the training set T h from the fingerprint feature vector set U, which contains r samples;

[0171] 5.3. Use the decision tree model proposed in step 4 to train the h-th round of iterative weak classifier model C h , as shown in the following formula:

[0172] C h =s_Dtree(T h )

[0173] 5.4. For the training set T h Each sample u in g , proceed to step 5.5 to step 5.7;

[0174] 5.5. When the hth round of iteration weak classifier model is used for sample u g The classification result C h (u g ) and sample u g The original category label v h If they are not the same, it proves that the classification fails. The error rate of the classifier model in the hth iteration is e h The calculation formula is as follows, w h (g) represents the training set T h In , the weight of the g-th sample is:

[0175] e h +=w h (g)

[0176] 5.6. In order to ensure that the classification of each weak classifier is better than that of a random classifier, the error rate e is required h Less than 0.5, when e h If it is greater than or equal to 0.5, the training ends and the process goes to step 5.2 to start a new cycle.

[0177] 5.7. When the error rate e is satisfied h If it is less than 0.5, then calculate α h , the formula is as follows:

[0178]

[0179] 5.8. When the hth round of iteration weak classifier model is used for sample u g The classification result C h (u g ) and sample u g The original category label v g If they are different, in order to ensure that the new round of weights calculated by the weight calculation formula conforms to the probability distribution, the standard factor Z is calculated. h , the formula is as follows:

[0180]

[0181] The sample u used in the new round of iteration g The weight value w h+1 (g) The calculation formula is as follows:

[0182]

[0183] 5.9. When the weak classifier model in the h-th round of iteration classifies the sample u g The classification result C h (u g ) is the same as the original class label v of the sample u g , it proves that the classification is successful, and the Z g formula is as follows: h The formula is as follows:

[0184]

[0185] w h+1 (g) The calculation formula is as follows:

[0186]

[0187] 5.10. After all training loops are completed, return the strong classifier f(x), and the formula is as follows:

[0188] f(x) = ∑α h C h .

[0189] Step 6. Use the constructed strong classifier to identify the collected data set.

[0190] In order to verify that the networked industrial control asset identification method based on the improved decision tree proposed by the present invention does have noise reduction ability, the traffic data detected by the networked industrial control asset identification method based on the improved decision tree of the present invention and the traffic data detected by the traditional detection method are respectively used as the data set after being selected and processed according to the fingerprint features of the present invention, and input into the C4.5 decision tree model, and the accuracy, precision, recall rate and F1 value are used as the evaluation criteria for the classification effect. The experimental results are as Figure 4 shown.

[0191] It can be seen from the experimental results that when using the data obtained by the detection model of the present invention, in the C4.5 decision tree classification, whether it is for the operating system or the device model, it has higher accuracy, precision and better coverage. This shows that the detection model of the present invention can effectively improve the data purity and reduce the noise impact in the detection of the operating system and device model of industrial control devices.

[0192] When the fingerprint features proposed by the present invention are used to construct a dataset and input into an unimproved decision tree model, the accuracy rate of the decision tree model exceeds 90%. Therefore, the fingerprint feature combination proposed by the present invention does have a good recognition and classification effect in the identification of industrial control device operating systems and device models.

[0193] In order to reduce the influence on the comparative experiment caused by the imbalance of the dataset between models during sampling testing due to uneven data distribution, the present invention uses the method of ten-fold cross-validation for the comparative experiment. The ten-fold validator will divide the dataset into ten parts, and conduct ten classification trainings and validations for each classifier. Each time an experiment is carried out, one part of the dataset is used as the validation set, and the other nine parts are used as the training set. Each time an experiment is carried out, the average value of each item of each category is taken as the result of one experiment. The average value is divided into weighted average and equal-weight average. The weighted average attaches importance to the category data with relatively large data quantity, and the equal-weight average sets the weights of all category data to one. Since the data category classification of the present invention is not average, the equal-weight average formula is selected for the average value calculation of the present invention. The ten-fold validator of each classifier finally takes the average value of the results of ten experiments as the classification performance of the classifier. The comparison results are as Figure 5 shown. It can be seen from the performance of each method in the identification of the operating system of networked industrial control devices that the performance of Naive Bayes is the worst in all four values. The performance of SVM in terms of accuracy and precision is similar to that of the decision tree and the method of the present invention. However, the performance of the SVM method in terms of recall rate and F1 value is significantly inferior to that of the decision tree and the method of the present invention. Compared with the traditional decision tree, the method for identifying networked industrial control assets based on the improved decision tree of the present invention has a slight increase of more than 1% in both accuracy and precision, and the performance in terms of recall rate and F1 value is also significantly better than that of the traditional decision tree model.

[0194] As Figure 6 , it can be seen from the comparative experiment on the identification of networked industrial control device models that since the classification labels reach 48, the classification at this time is a multi-class classification problem. Therefore, the performance of the SVM method in terms of accuracy and precision drops sharply, while the performance of the Naive Bayes model has a slight increase in the problem of device model identification. However, since the Naive Bayes model assumes complete independence between all attributes and does not care about the correlation between attributes, there is still a large gap in the overall performance compared with the decision tree and the method of the present invention. This also reflects indirectly that the industrial control network traffic features extracted by the invention are correlated with each other. Comparing the traditional decision tree and the method of the present invention, the method of the present invention has an improvement of more than 2% in both accuracy and precision, and the improvement in terms of recall rate and F1 value is more obvious. This shows that in the problem of identifying networked industrial control device models, the method of the present invention has obvious advantages.

[0195] In summary, the method for identifying the assets of networked industrial control devices proposed by the present invention has good performance in identifying industrial control operating systems and industrial control device models, and has certain advantages in classification and comparison with other traditional classification models in the identification of networked industrial control assets.

Claims

1. An Internet-connected industrial control asset identification method based on an improved decision tree, characterized in that, It includes the following steps; Step 1: Use the networked industrial control asset detection model to detect data; the detected information forms an industrial control device traffic data set; Step 2: Propose a fingerprint feature combination, combine the industrial control device traffic data set with this fingerprint feature combination, and perform data preprocessing through normalization or one-hot encoding to form a data set; Obtain data labels according to the detected information, and divide the data set into a networked industrial control device operating system asset data set and a networked industrial control device model asset data set according to the data labels; Step 3: According to the different importance of different features of the industrial control device traffic data set for the asset identification and classification problem, and at the same time to solve the differences between the sample data fingerprint features in the networked industrial control device operating system asset data set and the networked industrial control device model asset data set, calculate the feature weights; Step 4: Build a decision tree model based on the corrected feature weights, and use the product of the feature weights calculated in Step 3 and the Gini coefficient as the node splitting criterion to establish a decision tree model; consider the importance of the fingerprint features and the accuracy of the model after classifying the fingerprint features when establishing the decision tree model; Step 5: Combine the decision tree model built in Step 4 to establish a decision tree method based on Adaboost to improve the accuracy and reduce the coupling; 2. The method for identifying networked industrial control assets based on an improved decision tree according to claim 1, wherein The specific working process of the networked industrial control asset detection model in Step 1 is as follows; 1.1 Input the target host IP address; 1.2 Send an ICMP protocol data packet. When the target host responds with an Echo reply message, go to Step 1.3; If no response from the target host is received within the specified time, this round of detection ends, and Step 1.1 is restarted to input a new target host IP address; 1.3 Send a SYN data frame to the target port of the target host. When the target host replies with a SYN ACK data frame, perform Step 1.4; If the target host does not respond, this round of detection ends, and Step 1.1 is restarted to input a new target host IP address; 1.4 Select an industrial control protocol data packet according to the target host port opening situation and send it, and monitor the target host traffic; When the reply packet from the target host conforms to the industrial control protocol standard, perform Step 1.5; when the target host does not respond or responds incorrectly, this round of detection ends, and Step 1.1 is restarted to input a new target host IP address; 1.5 Check the honeypot data features of the target host traffic data, including the target host service provider, multi-port opening situation, and response time length. When the target host does not have the above honeypot data features, perform Step 1.

6. When the target host has any of the above honeypot data features, perform Step 1.9; 1.6 Use the shodan search engine interface to check the target host information. The check situations are divided into three types; when shodan marks the target host as a suspected honeypot, perform Step 1.

7. When shodan marks the target host as a confirmed honeypot, perform Step 1.

9. When it does not meet the above two situations, the target host is considered an industrial control device, and perform Step 1.10; 1.7 Wait for a fixed time, then resend the industrial control protocol connection request and update the industrial control protocol function command code; 1.

8. Check the communication behavior of the target host. When there is a behavior of rejecting connections or not supporting new industrial control protocol function command codes, it is considered that the target host has the behavior characteristics of a honeypot and proceed to step 1.9; otherwise, proceed to step 1.10; 1.

9. Mark the target host as a honeypot device, discard the traffic data of the target host, and end the detection; 1.

10. Mark the target host as an industrial control device, save the traffic data of the target host, and store the relevant detection information in the industrial control device traffic dataset.

3. The method for identifying networked industrial control assets based on an improved decision tree according to claim 2, wherein The specific content of step two is as follows: The fingerprint feature combination includes the conventional fingerprint features in the TCP / IP protocol data packet and the fingerprint features selected according to the characteristics of the industrial control device itself; among the mean, variance, range, maximum value, and minimum value of the variable data in the industrial control device traffic dataset, the fields with regular changes are subjected to data reduction and data aggregation to describe the periodic changes of the data flow characteristics; therefore, the fingerprint features selected according to the characteristics of the industrial control device itself include the dimension num, mean av, variance var, maximum value max, minimum value min, extreme value mr, port port, industrial control protocol service, whether to use industrial Ethernet ifet, longitude longitude, and latitude latitude; Combine the industrial control device traffic dataset with the fingerprint feature combination to obtain the combined dataset; 2.

1. Calculate the variance, mean, and range of the numerical fingerprint features in the combined dataset; Set the fingerprint feature A in the combined dataset i There are n values in total, where a ij is the j-th value of the fingerprint feature A i ; the mean value av(A i ) in the fingerprint feature A i is defined as follows: Fingerprint feature A i The variance var(A i ) is defined as follows: Fingerprint feature A i The range mr(A i ) is defined as follows: mr(A i ) = max(A i ) - min(A i ) 2.

2. Normalize the numerical fingerprint features in the combined dataset that do not conform to the standard normal distribution so that the overall mathematical distribution of this type of numerical fingerprint features conforms to the standard normal distribution; use Z-score normalization, and the normalization conversion process is shown in the following formula, where A is the fingerprint feature set and a i is the specific fingerprint feature value: 2.

3. Perform one-hot encoding on the discrete classification features in the combined dataset that cannot be infinitely subdivided to make the fingerprint feature values of this type ordered and continuous; 2.

4. According to the detection information collected in step 1.10, use regular expressions to extract the valid information and use it as the data label. According to the types of data labels, divide the dataset processed in steps 2.1 - 2.3 into the operating system asset dataset of the networked industrial control device and the model asset dataset of the networked industrial control device.

4. The method for identifying networked industrial control assets based on an improved decision tree according to claim 3, characterized in that The specific operations for the operating system asset dataset of the networked industrial control device and the model asset dataset of the networked industrial control device in step three are as follows: 3.

1. Randomly select a sample R from the operating system asset dataset of networked industrial control devices or the model asset dataset of networked industrial control devices l , where l is required to be less than or equal to the input random sampling number O. For each sample R l , perform the operations from step 3.2 to step 3.3; 3.

2. Calculate the sample R l and the Euclidean distance of each sample with the same category as R l and find the set H of k nearest neighboring samples of the same category. The Euclidean distance is calculated using the following formula, where x α represents the coordinates of the sample point P, and y α represents the coordinates of the sample point B; P and B are respectively the sample points in the operating system asset dataset of the networked industrial control device or the model asset dataset of the networked industrial control device; c is the dimension of the data point; 3.

3. Calculate the sample R l and the Euclidean distance of each sample that is different from R l to also find a set M of k nearest neighbor samples of different classes; 3.

4. Traverse the sample R l and its set of similar samples H. For each H in the set of similar samples H p perform the operation in step 3.5; 3.

5. When the fingerprint feature A i has different values in the sample R l and the sample H p add 1 to diff_rh, where diff_rh represents the intermediate value when calculating the feature weights of similar samples; 3.

6. Calculate the feature weight value w_same of the fingerprint feature A in the set H of similar samples of the same category as the sample R according to the following formula l where m is the number of iterations i in the set H of similar samples of the same category 3.

7. Traverse and sample R l For the category set CM in the heterogeneous sample set M of q Perform operations from step 3.8 to step 3.10 for each category C in CM; 3.

8. Traverse the sample R l in the outlier sample set M of q the sample set M(C q ) of class C, for each sample M q in M(C x ), perform the operation in step 3.9; 3.

9. When the fingerprint feature A i has different values on the sample R l and the sample M x then diff_rm is incremented by 1. diff_rm represents an intermediate value when calculating the feature weights of heterogeneous samples; 3.

10. Calculate the sample R according to the following formula l of the fingerprint feature A in the outlier sample set M i of the weight value w_diff, where s is the number of samples in M(C q ), class(R) represents the class of sample R, and P represents probability: w_diff+=P(C q )*diff_rm / (1 - P(class(R))*s*k) 3.

11. Calculate the feature weight W(A i ) of each fingerprint feature A in fingerprint feature set A i : i i )​ W(A i ) = W(A i ) - w_same + w_diff。 5. The method for identifying networked industrial control assets based on an improved decision tree according to claim 4, characterized in that The specific content of step four is as follows: 4.

1. Initialize and create a root node that includes the entire operating system asset dataset of the networked industrial control device or the model asset dataset of the networked industrial control device; 4.

2. Multiply each feature weight \(W(A)\) in the feature weight set \(W(A)\) calculated in Step 3 by the Gini coefficient \(Gini(D)\) to obtain the corrected Gini coefficient, which serves as the basis for selecting the feature for node splitting; i ) The Gini coefficient Gini(D) of the fingerprint feature A for the operating system asset dataset of the networked industrial control device or the model asset dataset of the networked industrial control device is calculated by the following formula; Q1 = {x ∈ D|A(x) = a} Q2 = D - Q1 Where Q1 is the sample set in the networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset where the fingerprint feature A takes the value a; Q2 is the sample set in the networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset where the fingerprint feature A does not take the value a; assume the networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset D i The class set E on the fingerprint feature A is divided into {E1, E2, ……, E n}; 4.

3. Determine the selected fingerprint feature A i as the node division criterion; the networked industrial control device operating system asset dataset or the networked industrial control device model asset dataset is divided into several subsets D i according to different values of the fingerprint feature A i , and delete the fingerprint feature A i from the fingerprint feature set A; 4.

4. For subset D i in each subset D ij , perform operations from step 4.5 to step 4.7; 4.

5. When D ij All instances in it belong to the same class, create a leaf node and label the node with the class; 4.

6. When D ij is empty, return to step 4.4; 4.

7. When the number of nodes in D ij is less than the number of neighboring samples k + 1 and the number of subset nodes is less than the random sampling number f + 1, create a leaf node and label the class with the largest number in D ij as the node class; 4.

8. When D ij does not meet all of the above conditions, go to step 4.2; 4.

9. Output the constructed decision tree model based on feature weight correction.

6. The method for identifying networked industrial control assets based on an improved decision tree according to claim 5, wherein The specific content of step five is as follows: 5.

1. Define a set of fingerprint feature vectors U = {u1, u2, u3, …, u t} for the operating system asset dataset of networked industrial control devices or the asset dataset of networked industrial control device models. The corresponding classification label set of U is defined as V = {v1, v2, v3, …, v t}. It is specified that the weights of each sample u in the feature vector set are the same, which are i.e., the initial feature weights Input the number of training loops N. Steps 5.2 to 5.9 are executed in each loop; 5.2 Weight coefficient α of the classifier model in the h-th round of iteration h The initial value is set to 0, α h represents the weight of the weak classifier in the h-th round of iteration in the final strong classifier. A training set T is extracted from the fingerprint feature vector set U h , which contains r samples; 5.

3. Use the decision tree model based on feature weight correction proposed in Step 4 to train the weak classifier model \(C\) in the \(h\)-th round of iteration h , which is represented by the following formula: C h = s_Dtree(T h ) 5.

4. For each sample u in the training set T h perform the operations in steps 5.5 to 5.7; g ​ 5.

5. When the classification result C g of the h-th round of iterative weak classifier model for the sample u h (u g ) is different from the original class label v g of the sample u h , it proves that the classification by typing fails. The error rate e h of the h-th round of iterative classifier model is calculated as follows, where w h (g) represents the weight of the g-th sample in the training set T h : e h += w h (g) 5.

6. To ensure that the classification performance of each weak classifier is better than that of a random classifier, it is required that the error rate e h is less than the set threshold. When e h is greater than or equal to the set threshold, end this training, go to step 5.2, and start a new round of loop; 5.

7. When the error rate e h is less than the set threshold, then calculate α h , and the formula is as follows: 5.

8. When the classification result C g of the h-th round of iterative weak classifier model for the sample u h (u g ) is different from the original class label v g of the sample u g , to ensure that the new round of weights calculated by the weight calculation formula conform to the probability distribution, calculate the normalization factor Z h . The formula is as follows: The sample u used in the new round of iteration g The weight value w h+1 (g) The calculation formula is as follows: 5.

9. When the classification result C g of the h-th round of iterative weak classifier model for the sample u h (u g ) is the same as the original class label v g of the sample u g , it is proved that the classification by typing is successful. The Z h formula is as follows: w h+1 (g) The calculation formula is as follows: 5.

10. After all training loops end, return the strong classifier f(x), and the formula is as follows: f(x) = ∑α h C h 。 7. The method for identifying networked industrial control assets based on an improved decision tree according to claim 6, wherein The threshold is set to 0.5 in step 5.

Citation Information

Patent Citations

  • Passive industrial control system asset identification method and device

    CN114584497A

  • Equipment type fingerprint generation method, identification method, equipment and medium

    CN115589362A