Online industrial control intrusion detection system considering rule reduction
By designing an online industrial control intrusion detection system, using technical means such as data collection, evaluation system construction, and rule simplification, the problem of industrial control systems decreasing intrusion detection accuracy and explosion of rules combination in dynamic environments is solved, and efficient and accurate intrusion detection is achieved.
Patent Information
- Application Number
- CN202411943953.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-16
AI Technical Summary
The existing industrial control system intrusion detection methods are difficult to maintain good performance in dynamic environments, especially when data distribution and quantity change, resulting in a problem of degradation of detection accuracy and explosive rule combination.
An online industrial control intrusion detection system that considers rule reduction is designed. Through data collection, construction of evaluation systems, calculation of confidence distribution, rule reduction and dynamic rule updates, real-time intrusion detection and rule optimization of industrial control systems are realized.
By building a comprehensive intrusion detection index system, integration of expert systems, intelligent reasoning, and application of rule reduction algorithms, the system improves the inference efficiency and accuracy of the model, and can maintain a high intrusion detection accuracy in a dynamic environment.
Smart Images

Figure CN120012073A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of fault detection, and in particular relates to an online industrial control intrusion detection system considering rule simplification. Background Art
[0002] It is of great significance to simplify the rules for online fault diagnosis of industrial control systems. Industrial control systems are the core of modern industrial production and are directly related to production efficiency, safe operation, and the stable operation of national critical infrastructure. However, the industrial control system does not fully consider network security during design and security lags caused by long-term use, which increases the risk of intrusion. Because the environment in which the industrial system is located and its own health status are constantly changing, if it is not updated online, the accuracy of fault diagnosis will be reduced. Once an intrusion problem occurs in the industrial system, it will lead to a series of major problems such as abnormal system operation, information leakage, and economic benefit loss. Therefore, it is of great value to study intrusion detection of industrial systems.
[0003] In the current research on intrusion detection methods, representative methods include Long Short-Term Memory (LSTM), Support Vector Machine (SVM), Convolutional Neural Network (CNN) and Belief Rule Base (BRB). The LSTM method focuses on processing the long-term dependency of sequence data. Unlike ordinary neural networks, LSTM has structures such as forget gate, input gate and output gate, which can effectively maintain and forget the information in the sequence through the gating mechanism, thereby solving the gradient vanishing problem in long sequence data. However, as the indicator system expands, the model training time increases, resulting in poor timeliness. For the SVM method, when processing large-scale data, the training time and space complexity are high and it is not easy to expand. At the same time, the choice of kernel function for SVM depends more on the characteristics of the actual problem, and certain tuning work is required, resulting in poor timeliness. The CNN method achieves autonomous learning by simulating human neural processes and has strong fitting ability, but it is a black box model with strong dependence on large samples and insufficient credibility of the results. In contrast, the BRB method is more intuitive and effective in equipment fault diagnosis due to its simple and efficient evaluation process. This method has been widely used in medical decision-making, risk analysis, pattern recognition, industrial control intrusion detection and other fields.
[0004] As an extension of the BRB method, the Extended belief rule base (EBRB) solves the combinatorial rule explosion problem by expanding the antecedents of each sample. This model not only inherits the efficiency of the fuzzy rule base system modeling method, but also uses the belief rule representation framework in the belief rule base system that represents multiple uncertain information through belief distribution. Ye proposed a method of using EBRB to predict environmental governance costs, which effectively reduced the complexity of the model through dimensionality reduction and achieved good results. Yang proposed a joint optimization EBRB for bridge risk prediction, which optimized the parameters and structure of the model and had good accuracy.
[0005] Although the above model performs well in intrusion detection in a static environment, it is difficult for the above method to continue to maintain good performance in a dynamic environment over time; in a dynamic environment, the distribution and quantity of data often change, and the performance of the industrial control detection system may be seriously affected; although EBRB has the ability to process semi-quantitative information and has high performance and high interpretability, due to the complexity of the industrial system, there are too many input attribute values and the equipment runs for too long, and it takes a long time to update the data to maintain the accuracy of intrusion detection, which will lead to the problem of rule combination explosion; how to timely and accurately identify intrusion detection categories in a dynamic environment and reasonably simplify rules has become a problem that needs to be solved urgently.
[0006] Based on this, the present invention designs an online industrial control intrusion detection system considering rule simplification to solve the above problems. Summary of the invention
[0007] The purpose of the present invention is to solve the problems in the above-mentioned background technology and to propose an online industrial control intrusion detection system taking rule simplification into consideration.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] An online industrial control intrusion detection system considering rule simplification includes:
[0010] A data collection unit, which obtains a measured data set for natural gas pipeline status intrusion detection;
[0011] Construct an evaluation system module, build an index evaluation system for natural gas pipeline status intrusion detection, and optimize model parameters through optimization algorithms;
[0012] A calculation unit, calculating the confidence distribution of each indicator;
[0013] The simplification module simplifies the rules in the extended confidence rule base obtained in the calculation unit using the FCM algorithm;
[0014] The data update module, after acquiring new update data, uses a domain value-based rule update strategy to update the rule base to obtain a simplified rule base;
[0015] The testing module uses the updated model to predict the test set, studies the predicted results based on matching and inference, and compares them with the traditional model.
[0016] As a further description of the above technical solution: In the data collection unit, the five features with the highest feature importance are selected as the final features, and one normal flow and four attack flows simulating intrusion of the natural gas pipeline status under normal operating conditions are collected. The four attack flow types are DOS, CI, RI and Recon.
[0017] As a further description of the above technical solution: In the module for constructing the evaluation system, it is involved to determine the key indicators that affect the intrusion of the natural gas pipeline status, and use the reference values of the online evaluation of the FCM algorithm based on expert knowledge (FBE) to adjust the parameters of the premise attributes to achieve the optimal diagnostic performance.
[0018] As a further description of the above technical solution: In the calculation unit, the measured value of each indicator is evaluated using a confidence function to generate a confidence distribution, which expresses the confidence that each indicator belongs to a normal or abnormal state under a given measured value.
[0019] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0020] 1. The present invention adopts a comprehensive intrusion detection index system: through in-depth mechanism analysis, the present invention constructs a comprehensive industrial control system intrusion detection index system. This system is based on the working principle and intrusion characteristics of the equipment, and accurately defines a series of key indicators of intrusion traffic for detection and evaluation, ensuring that the equipment status can be fully monitored from multiple intrusion dimensions; expert system integration and intelligent reasoning: the present invention uses the knowledge of industry experts to establish an online evaluation model that uses expert knowledge for reference values. The model can perform joint reasoning on various performance indicators. This expert knowledge-based method not only improves the accuracy of diagnosis, but also enhances the model's ability to identify complex intrusion types, especially when facing large fluctuations in the environment where the equipment is located and facing new types of intrusions, it can make more reasonable decisions; Application of rule reduction algorithm: In order to further improve the reasoning efficiency and accuracy of the model, the present invention adopts an advanced rule reduction algorithm to adjust and optimize the model parameters. This method can achieve efficient simplification of rules, and can maintain good detection accuracy in a simulated online environment while reducing the number of rule bases by half; Efficient dynamic environment adaptability: The present invention constructs an online industrial control intrusion detection model, so that the model continuously updates the type of intrusion traffic in a changing dynamic environment, so that the model can achieve a higher intrusion detection accuracy; In general, the present invention provides an efficient, accurate, concise and reliable intrusion detection rule simplification method for intrusion detection of industrial control systems, which is particularly suitable for complex systems with a large number of rules and intrusion detection systems with long running times. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A modeling diagram of an online industrial control intrusion detection system based on an online updated EBRB model in consideration of rule simplification proposed by the present invention;
[0022] Figure 2 A new reference value online evaluation method diagram of EBRB of an online industrial control intrusion detection system considering rule simplification proposed by the present invention;
[0023] Figure 3 A diagram of an EBRB rule simplification method for an online industrial control intrusion detection system considering rule simplification proposed by the present invention;
[0024] Figure 4 This is an EBRB reasoning mechanism diagram of an online industrial control intrusion detection system considering rule simplification proposed by the present invention;
[0025] Figure 5 This is a diagram of the EBRB rule update process of an online industrial control intrusion detection system considering rule simplification proposed by the present invention;
[0026] Figure 6This is a model test result diagram of an online industrial control intrusion detection system that considers rule simplification proposed by the present invention. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] Please see attached Figure 1 -Attached Figure 6 The present invention provides a technical solution: an online industrial control intrusion detection system considering rule simplification, comprising:
[0029] A data collection unit, which obtains measured data of the natural gas pipeline status;
[0030] By measuring various indicators of the natural gas pipeline status, the measured data set is obtained;
[0031] Construct an evaluation system module, build an indicator evaluation system for natural gas pipeline status, and optimize model parameters through optimization algorithms;
[0032] The natural gas pipeline status data is first collected in real time by sensors in the pipeline system, and then the data packets are sent to the controller through the Modbus industrial data transmission protocol communication network. The controller calculates and generates control command data packets based on the received status data, and finally sends them to the actuator through the network for regulation. For the intrusion detection of the natural gas pipeline status, some performance indicators can usually be used to quantitatively or qualitatively characterize its status. When constructing the intrusion detection index system of the natural gas pipeline status, the basic principles of scientificity, systematicity, comprehensiveness, and dynamicity should be followed. On the one hand, starting from the characteristics of the natural gas pipeline status data, data with higher eigenvalues can be selected as the final features of the experiment, because the information contained in the features with low eigenvalues cannot effectively distinguish the attack categories. This experiment selects the five features with the highest feature importance as the final features, and uses the FCM algorithm based on expert knowledge to evaluate the reference value online for modeling. As the premise attribute of the model, the model selects 5 different traffic types to simulate intrusion, including one normal traffic and four attack traffic. Then the reference value of the premise attribute is optimized using the FCM algorithm based on expert knowledge. The optimization process is as follows Figure 2 shown.
[0033] The computing unit converts data into rules;
[0034] Then, for each piece of data in the old data, it is necessary to convert it into a rule. The specific process is shown in formulas (1)-(4); one of the rules can be expressed as:
[0035] If x1 is{(A 1,1 ,α 1,1 ),(A 1,2 ,α 1,2 ),(A 1,3 ,α 1,3 ),(A 1,4 ,α 1,4 ),(A 1,5 ,α 1,5 )}
[0036] ∧x2 is{(A 2,1 ,α 2,1 ),(A 2,2 ,α 2,2 ),(A 2,3 ,α 2,3 ),(A 2,4 ,α 2,4 ),(A 2,5 ,α 2,5 )}
[0037] ∧x3 is{(A 3,1 ,α 3,1 ),(A 3,2 ,α 3,2 ),(A 3,3 ,α 3,3 ),(A 3,4 ,α 3,4 ),(A 3,5 ,α 3,5 )}
[0038] ∧x4 is{(A 4,1 ,α 4,1 ),(A 4,2 ,α 4,2 ),(A 4,3 ,α 4,3 ),(A 4,4 ,α 4,4 ),(A 4,5 ,α 4,5 )}
[0039] ∧x5 is{(A 5,1 ,α 5,1 ),(A 5,2 ,α 5,2 ),(A 5,3 ,α 5,3 ),(A 5,4 ,α 5,4 ),(A5,5 ,α 5,5 )}
[0040] Then D is{(D1,β1),(D2,β2),(D3,β3),(D4,β5),(D5,β5)}
[0041] The extended confidence rule can be expressed as follows:
[0042]
[0043] with attribute weight{δ1,...,δ M}
[0044] and rule weightθ k
[0045] (1) There are M premise attributes x i (i=1,...,M), N result attributes D n (n=1,...,N). For each premise attribute, there is J i (i=1,...,M) reference points A i,j (i=1,...,M;j=1,...,J i ), in each rule, each antecedent attribute has a corresponding confidence distribution for a given reference value. i,j , the corresponding confidence distribution is Corresponding to the confidence of each conclusion. δ1 represents the attribute weight, θ k Indicates the rule weight.
[0046] Through the utility-based information conversion method in BRB, each sample in the sample set can be converted into a distributed confidence distribution form, where the distributed confidence of the j-th premise attribute is as follows:
[0047]
[0048] in
[0049]
[0050] Similarly, the output value y can be k Expressed as the corresponding distributed confidence:
[0051]
[0052] Among them, R i represents the i-th extended confidence rule, represents the confidence distribution of the u-th input value of its k-th attribute, Represents the confidence distribution generated by it for the nth result.
[0053] Simplification module, which simplifies rules based on the FCM algorithm;
[0054] The method of rule simplification is as follows Figure 3 express:
[0055] Step 1: Input the training data into the EBRB model. Experts provide reference values, and EBRB converts the data into corresponding rules.
[0056] Step 2: Use the initial rules as the input of the FCM clustering algorithm. The number of cluster centers is given by the experts, and the FCM algorithm is used to divide each rule into the corresponding cluster.
[0057] Step 3: Merge the rules in the same cluster center to obtain simplified rules.
[0058] The data update module uses a domain value-based rule update strategy to update the rule base;
[0059] For newly observed samples, they are not directly converted into confidence rules and added to the old EBRB, but are selectively added to the confidence rule base. If the current model predicts the new sample incorrectly, it means that the model lacks the confidence distribution of the sample, and the sample should be added to the rule base. If the current model predicts the new sample correctly, it means that the model already has the confidence distribution related to the sample. However, if the sample is directly abandoned and added to the EBRB, the subtle changes in the confidence distribution will not be detected, which may affect the performance of the model. Therefore, it is necessary to merge the correctly predicted samples into the EBRB. The most basic idea of fusion is to find the confidence distribution in the rule base that is closest to the sample, and then merge the confidence distributions of the two. In order to find the confidence distribution closest to the sample, the domain partitioning idea proposed by Yang
[24] is used.
[0060] The test module uses a rule-based reduction strategy for modeling and analysis;
[0061] Table 1 Information about pipeline datasets
[0062]
[0063]
[0064]
[0065] As can be seen from the table, some features have an importance score of 0, which means that the information contained in these features cannot effectively distinguish the category. In fact, the variance of these features is 0, which means that their values remain unchanged, which makes them unable to be used to distinguish the attack category. Finally, this experiment selected the five features with the highest feature importance as the final features for modeling.
[0066] The dataset contains five different traffic types, including one normal traffic and four attack traffic types to simulate intrusion. The four attack traffic types are DOS, CI, RI and Recon. Command injection (CI) includes malicious state command injection (MSCI), malicious parameter command injection (MPCI), malicious function code command injection (MFCI), and Response injection (RI) is divided into two attack types, namely Malicious response injection (NMRI) and complex malicious response injection (CMRI). The following table specifically shows the number of different traffic in the dataset. In general, normal traffic occupies the majority of the dataset, and DOS attack type samples are relatively small.
[0067] Table 2 Distribution of intrusion samples in the dataset.
[0068]
[0069] Table 3 Utility values of the antecedent attributes optimized using the FCM algorithm based on expert knowledge.
[0070]
[0071] Table 4 Utility values of intrusion outcomes set by experts.
[0072]
[0073] After each training set is converted into the rules shown in step 4, it is input into the test set for inference.
[0074] The inference mechanism of the model is as follows Figure 4 As shown, it is divided into the following three steps:
[0075] Step 1: Calculate individual matching. Assume that a vector consisting of the premise attributes of the sample to be tested is represented as x(x i; i=1,...,M), according to formulas (2)-(4), it can be converted into distributed confidence, where the distributed confidence of the i-th attribute is expressed as follows:
[0076] S(x i )={(A i,j , α i,j ); j = 1, ..., J i} (5)
[0077] Then, the matching degree of the sample to be tested relative to each premise attribute in each EBR can be calculated by formulas (6)-(7):
[0078]
[0079] in It represents the distance between the confidence distribution of the i-th premise attribute in the given test data and the i-th premise attribute in the k-th rule in the rule base. It represents the matching degree between the i-th premise attribute in the given test data and the i-th premise attribute in the k-th rule in the rule base. It can be seen from formula (6) that if The larger the distance, the less the two attributes match. If the distance between two attributes exceeds 1, they are considered completely mismatched. To express.
[0080] Step 2: Calculate activation weights. According to step 1, calculate the matching degree of sample x for each attribute in each rule. Then use the rule weight θ k (k=1, ..., L) and attribute weight δ i (i=1, ..., M), the activation weight of each confidence rule EBR for sample x can be calculated, where the activation weight calculation formula of the kth confidence rule is as follows:
[0081]
[0082] Step 3: Synthesize activation rules. After calculating the activation rule weights, all rules with weights greater than 0 are expanded as activated new rules. Through the evidence reasoning parsing algorithm, these activated rules are merged into one rule as the final result. Without loss of generality, assume that L · If all the extended confidence rules are activated rules, the resulting confidence distribution of the synthesized rules is expressed as follows:
[0083]
[0084] in represents the confidence of the nth result in the kth rule, ω k represents the activation weight of the kth rule, β n Represents the combined confidence and satisfies β n ≥0,
[0085] The model update process is as follows Figure 5 As shown, there are three steps in total:
[0086] Step 1: Compare the newly observed samples with the rules in the rule base
[0087] Step 2: If the current model makes an error in predicting a new sample, it means that the model lacks the confidence distribution of the sample and the sample should be added to the rule base.
[0088] Step 3: If the current model predicts the new sample correctly, it means that the model already has the confidence distribution related to the sample. However, if the sample is directly abandoned and added to the EBRB, the subtle changes in the confidence distribution will not be detected, which may affect the performance of the model. Therefore, it is necessary to integrate the correctly predicted samples into the EBRB.
[0089] During the experiment, in order to construct the online update model and verify the model effect, the number of the second type of attacks in the data set was first artificially reduced, and the number of the second type of attacks in the test data set was increased at a specific time, so as to simulate the change of attack types in a dynamic environment. To this end, the data set is divided into three parts. The first part is used as the old data train_old to build the offline intrusion detection model. The second part is used as the newly observed data train_new to update the model online. The last part is used as the test data test to test the model and verify the effect of the model. The effect diagram of the model is shown in the figure below. Figure 6 The comparison of the effects of different models is shown in the following table:
[0090] Table 5 Comparison of the number of rules and model accuracy of different models
[0091]
[0092] The above table compares the number of rules and their accuracy of various confidence rule base models. It can be seen that for the BRB model, interval BRB model and EBRB model, OEBRB-r has good accuracy and can adapt well to the conversion of attack types in a dynamic environment. For OEBRB, OEBRB-r can effectively reduce the number of rules to about half of the original number without greatly reducing the accuracy, which proves the effectiveness of the model.
[0093] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. An online industrial control intrusion detection system considering rule simplification, characterized in that: include: A data collection unit, which obtains a measured data set for natural gas pipeline status intrusion detection; Construct an evaluation system module, build an index evaluation system for natural gas pipeline status intrusion detection, and optimize model parameters through optimization algorithms; A calculation unit, calculating the confidence distribution of each indicator; The simplification module simplifies the rules in the extended confidence rule base obtained in the calculation unit using the FCM algorithm; The data update module, after acquiring new update data, uses a domain value-based rule update strategy to update the rule base to obtain a simplified rule base; The testing module uses the updated model to predict the test set, studies the predicted results based on matching and inference, and compares them with the traditional model.
2. The online industrial control intrusion detection system considering rule simplification according to claim 1 is characterized in that: In the data collection unit, the five features with the highest feature importance are selected as the final features. One normal flow and four attack flows are collected to simulate intrusion when the natural gas pipeline is in normal operation. The four attack flow types are DOS, CI, RI and Recon.
3. The online industrial control intrusion detection system considering rule simplification according to claim 1 is characterized in that: The module for constructing the evaluation system involves determining key indicators that affect the intrusion of the natural gas pipeline status, and using reference values of the online evaluation of the FCM algorithm based on expert knowledge (FBE) to adjust the parameters of the premise attributes to achieve optimal diagnostic performance.
4. The online industrial control intrusion detection system considering rule simplification according to claim 1 is characterized in that: In the calculation unit, the measured value of each indicator is evaluated using a confidence function to generate a confidence distribution, which expresses the confidence that each indicator belongs to a normal or abnormal state under a given measured value.