An industrial control system anomaly type detection method, system, device, and medium
Patent Information
- Application Number
- CN202411025629.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-07-30
AI Technical Summary
但SVM会受到数据集大小的影响,由于内存不足或长时间的训练而导致系统故障
[0015]有益效果:本发明提供的工业控制系统异常类型检测方法,先对特征进行标准化及降维处理,以减小计算机运行压力,再利用多目标粒子群优化算法优化支持向量机的参数和选择特征,从而既得到支持向量机的最优参数,又得到与异常类型相关的特征子集;根据支持向量机的最优参数和特征子集得到的工控系统异常检测模型,运行效率高且类型预测的准确率高,能够对工业控制系统中不同类别的攻击进行准确分类。
Smart Images

Figure CN118709059B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of industrial control system safety detection technology, specifically relating to a method, system, device, and medium for detecting abnormal types in industrial control systems. Background Technology
[0002] In recent years, various innovative applications such as the Industrial Internet, intelligent manufacturing, and the Internet of Things have developed rapidly. The number of networked control devices and data exchange facilities in Industrial Control Systems (ICS) is increasing, and the data access methods are becoming more diverse. These changes make industrial control devices more vulnerable to various types of cyberattacks during operation.
[0003] Currently, intrusion detection systems (IDS) employing Support Vector Machines (SVMs) can monitor network data transmission and detect abnormal behavior through data analysis. However, SVMs are affected by dataset size, and system failures can occur due to insufficient memory or long training times. In real-world scenarios, data often contains a large number of irrelevant or outdated features, and SVMs lack the ability to immediately grasp feature importance, which significantly increases computational complexity and may lead to weaker prediction accuracy. Summary of the Invention
[0004] To address the problems raised in the background art, the present invention provides a method, system, device, and medium for detecting anomaly types in industrial control systems.
[0005] The technical solution of the present invention is as follows: This invention provides a method for detecting anomaly types in industrial control systems, comprising the following steps: S1: Acquire industrial control system operation data, including: tags, command addresses, response addresses, command memory counts, communication read functions, response write functions, sub-functions, setpoints, control modes, control schemes, and measured values. Calculate the standard deviation of features in the industrial control system operation data, delete features with a standard deviation of zero, scale the features to have a mean of zero and a variance of 1, perform dimensionality reduction processing, and construct a feature dataset. S2: Based on the feature dataset, the optimal parameters of the feature subset and the support vector machine are obtained through a multi-objective particle swarm optimization algorithm, specifically: S21: Initialize the parameters of the multi-objective particle swarm optimization algorithm. The parameters include the position, velocity, mutation rate exponent, and preset number of iterations of the particles. The position of each particle in the particle swarm is a combination of the parameters of the support vector machine and the features selected from the feature dataset. Different particle positions have different combinations of parameters and different combinations of features. S22: Construct a fitness value vector based on classification accuracy and feature scores, and calculate the particle's fitness value vector based on the particle's position; S23: If the particle's fitness value vector is better than the particle's historical best fitness value vector, then update the particle's best position according to the particle's fitness value vector. If a particle's fitness value vector is not better than its historical best fitness value vector, then the particle's best position is not updated. S24: Determine the non-dominated solution based on the particle's optimal position; If the repository is not full, add non-dominated solutions directly to the repository; If the repository is full, calculate the crowding degree of non-dominated solutions in the repository, remove the non-dominated solution with the highest crowding degree, and add the new non-dominated solution to the repository. S25: The position of the particle with the least crowded non-dominated solution in the repository is taken as the optimal position of the particle swarm; S26: Based on the particle's velocity, optimal particle position, and optimal particle swarm position in the current round, update the particle's velocity to obtain the updated particle velocity; Update the particle's position based on the updated particle's velocity and the particle's position in the current round to obtain the particle's first position; S27: Based on the particle's first position, mutation rate exponent, preset iteration number, and current round, obtain the mutation probability, and generate a random number based on the mutation probability; If the random number is less than the mutation probability, then the first position of the particle is mutated to obtain the updated position of the particle. If the random number is not less than the mutation probability, then no mutation operation is performed, and the first position of the particle is used as the position of the updated particle. S28: If the preset conditions are met, the optimal parameters of the feature subset and the support vector machine are obtained based on the optimal position of the particle swarm. If the preset conditions are not met, execute S22 until the preset conditions are met, and obtain the feature subset and the optimal parameters of the support vector machine based on the optimal position of the particle swarm. S3: Match the optimal parameters to the support vector machine, construct a binary tree support vector machine based on binary tree theory, and obtain the anomaly detection model of the industrial control system after training with feature subsets; S4: Acquire the operating data of the industrial control system to be tested, and after the industrial control system anomaly detection model performs the test, output the type detection result.
[0006] In step S22, a fitness vector is constructed based on classification accuracy and feature scores. Based on the particle's position, the particle's fitness value vector is calculated, specifically as follows: According to the formula: Calculate the fitness value vector of the particles; In the formula, Let be the particle's fitness value vector. For classification accuracy, These are characteristic scores.
[0007] The classification accuracy is based on the formula: Calculate the classification accuracy; In the formula, The number of samples that are true negatives and correctly predicted as negative. The number of samples that are true positives and correctly predicted as positive. The number of positive samples that were incorrectly predicted as negative due to false negatives; The number of negative samples that were incorrectly predicted as positive (false positives). The feature score is calculated according to the formula: Calculate the characteristic scores; In the formula, The number of features selected from the feature dataset. This represents the total number of features in the feature dataset.
[0008] In step S26, the particle velocity is updated based on the particle velocity, optimal particle position, and optimal particle swarm position in the current round, resulting in the updated particle velocity. Specifically: According to the formula: Update the particle velocity to obtain the updated particle velocity; In the formula, k is the iteration round; Let be the velocity of the i-th particle in the (k+1)-th round; Let be the velocity of the i-th particle in the k-th round; Let be the optimal position of the i-th particle; Let i be the position of the i-th particle in the k-th round; Inertial weights are used to control the inertia of particle movement. The first cognitive coefficient represents the weight of the particle's own experience; The second cognitive coefficient represents the weight of the particle swarm's experience; and All are random numbers between 0 and 1; This is the optimal position for the particle swarm.
[0009] In step S26, the particle's position is updated based on the updated particle velocity and the particle's position in the current round to obtain the particle's first position, specifically as follows: According to the formula: Update the particle's position; In the formula, Let i be the position of the i-th particle in the (k+1)th round; Let i be the position of the i-th particle in the k-th round; Let be the velocity of the i-th particle in the (k+1)-th round.
[0010] In step S27, the mutation probability is obtained based on the particle's first position, mutation rate exponent, preset iteration number, and current round, according to the formula: , thus obtaining the mutation probability; In the formula, The mutation probability, It is the variability index. To preset the number of iterations, This refers to the current round.
[0011] In S1, the features are scaled to have a mean of zero and a variance of 1, specifically as follows: According to the formula: , which scales the features to have a mean of zero and a variance of 1; In the formula, These are the standardized eigenvalues. These are the original eigenvalues; Original eigenvalues The mean; Original eigenvalues The standard deviation.
[0012] The present invention also provides an industrial control system anomaly type detection system, comprising: Feature dataset construction module: used to acquire industrial control system operation data, including: tags, command addresses, response addresses, command memory counts, communication read functions, response write functions, sub-functions, setpoints, control modes, control schemes, and measured values. It calculates the standard deviation of features in the industrial control system operation data, deletes features with a standard deviation of zero, scales the features to have a mean of zero and a variance of 1, performs dimensionality reduction processing, and constructs the feature dataset. Optimization module: Based on the feature dataset, the optimal parameters of the feature subset and support vector machine are obtained through a multi-objective particle swarm optimization algorithm, specifically: The parameters of the multi-objective particle swarm optimization algorithm are initialized. The parameters include the position, velocity, mutation rate exponent, and preset number of iterations of the particles. The position of each particle in the particle swarm is a combination of the parameters of the support vector machine and the features selected from the feature dataset. Different particle positions have different parameter combinations and different feature combinations. A fitness value vector is constructed based on classification accuracy and feature scores, and the fitness value vector of a particle is calculated based on the particle's position. If a particle's fitness value vector is better than its historical best fitness value vector, then the particle's optimal position is updated based on the particle's fitness value vector. If a particle's fitness value vector is not better than its historical best fitness value vector, then the particle's best position is not updated. Determine the non-dominated solution based on the particle's optimal position; If the repository is not full, add non-dominated solutions directly to the repository; If the repository is full, calculate the crowding degree of non-dominated solutions in the repository, remove the non-dominated solution with the highest crowding degree, and add the new non-dominated solution to the repository. The position of the particle with the least crowded non-dominated solution in the repository is taken as the optimal position of the particle swarm. Based on the particle's velocity, optimal particle position, and optimal particle swarm position in the current round, update the particle's velocity to obtain the updated particle velocity. Update the particle's position based on the updated particle's velocity and the particle's position in the current round to obtain the particle's first position; Based on the particle's first position, mutation rate exponent, preset iteration number, and current round, the mutation probability is obtained, and a random number is generated based on the mutation probability. If the random number is less than the mutation probability, then the first position of the particle is mutated to obtain the updated position of the particle. If the random number is not less than the mutation probability, then no mutation operation is performed, and the first position of the particle is used as the position of the updated particle. If the preset conditions are met, the optimal parameters of the feature subset and the support vector machine are obtained based on the optimal position of the particle swarm. If the preset conditions are not met, the fitness value vector of the particles is recalculated until the preset conditions are met. Based on the optimal position of the particle swarm, the optimal parameters of the feature subset and the support vector machine are obtained. The industrial control system anomaly detection model establishment module is used to match the optimal parameters to the support vector machine. Based on the binary tree theory, a binary tree support vector machine is constructed. After training with feature subsets, the industrial control system anomaly detection model is obtained. Type detection module: Used to acquire the operating data of the industrial control system to be detected, and after being detected by the industrial control system anomaly detection model, output the type detection result.
[0013] The present invention also provides an industrial control system anomaly detection device, including a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the industrial control system anomaly detection method.
[0014] The present invention also provides an industrial control system anomaly detection medium for storing a computer program, wherein the computer program, when executed by a processor, implements the industrial control system anomaly detection method.
[0015] Beneficial effects: The industrial control system anomaly detection method provided by this invention first standardizes and reduces the dimensionality of features to reduce the computer's operating load. Then, it uses a multi-objective particle swarm optimization algorithm to optimize the parameters of the support vector machine and select features, thereby obtaining both the optimal parameters of the support vector machine and the feature subset related to the anomaly type. The industrial control system anomaly detection model obtained based on the optimal parameters of the support vector machine and the feature subset has high operating efficiency and high accuracy in type prediction, and can accurately classify different types of attacks in the industrial control system. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating an anomaly detection method for an industrial control system in one embodiment. Detailed Implementation
[0017] The following examples are intended to illustrate the present invention, and not to further limit the invention.
[0018] This invention provides a method for detecting anomaly types in industrial control systems, such as... Figure 1 As shown, it includes the following steps: S1: Acquire industrial control system operation data, including: tags, command addresses, response addresses, command memory counts, communication reading functions, response reading functions, sub-functions, setpoints, control modes, control schemes, and measured values. Calculate the standard deviation of features in the industrial control system operation data, delete features with a standard deviation of zero, scale the features to have a mean of zero and a variance of 1, perform dimensionality reduction processing, and construct a feature dataset.
[0019] The command address represents the target address of the command. Different attacks may target different command addresses; by analyzing changes in the command address, abnormal command injection can be identified.
[0020] The response address represents the address from which the response originated. Malicious response injection attacks often forge response addresses; analyzing anomalies in the response address can help identify this type of attack.
[0021] Command memory count indicates the number of memory regions involved in a command. Anomalies may involve unusual memory region counts; monitoring this characteristic can help detect malicious memory operations.
[0022] The Comm Read Function represents the function code of a read command. Different types of read operations correspond to different function codes. Abnormal or malicious read operations will use different function codes, which can be used for identification.
[0023] The Resp Read Function indicates the function code used to respond to a read command. Malicious responses may use abnormal function codes; monitoring this characteristic can help identify anomalous responses.
[0024] A subfunction represents a sub-function code of a command. Malicious command injection attacks may use specific subfunction codes; analyzing the usage of subfunction codes can help identify malicious commands.
[0025] A setpoint represents the set value of a control system. An anomalous attack might modify the setpoint to compromise the system; monitoring changes to the setpoint can help detect such attacks.
[0026] Control mode refers to the operating mode of a control system. Different attacks may attempt to switch control modes; by monitoring changes in control mode, abnormal operations can be identified.
[0027] A control scheme represents the configuration of a control system. Attackers may attempt to modify the control scheme to disrupt the system; analyzing changes to the control scheme helps in detecting such attacks.
[0028] Measurement: Represents the actual measured value of the system. Attackers may forge or tamper with measurement values. By monitoring abnormal changes in measurement values, forged data injection attacks can be identified.
[0029] In addition, the operating data of the industrial control system also includes: command memory (memory locations associated with commands), response memory (memory locations associated with responses), response memory count (number of memory regions involved in the response), communication write function (function code for writing commands), response write function (function code for writing responses to commands), error status (current error status), error count (number of times errors have occurred), cyclic redundancy check (CRC), function code (function code), command (command content), response (response content), rate (transmission rate), pump (pump status), solenoid valve (solenoid valve status), CRC check rate, time (time stamp or time interval), and result (classification result).
[0030] To reduce the computer's workload, and considering that the standard deviation of some features is close to zero, indicating that the feature has no difference across samples and has no effect on classification, features with a standard deviation of zero are deleted, and according to the formula: , which scales the features to have a mean of zero and a variance of 1; In the formula, These are the standardized eigenvalues. These are the original eigenvalues; Original eigenvalues The mean; Original eigenvalues The standard deviation.
[0031] In addition, since the number of some types of samples in this dataset differs too much from that of normal samples, in order to avoid overfitting, the SMOTE algorithm is used to oversample the fewer types of samples to balance the dataset. For example, when the number of samples of a certain type is less than 5% of the normal samples, the number of samples of that type is increased to 5% of the normal samples and added to the dataset to synthesize new samples.
[0032] Regarding dimensionality reduction, linear discriminant analysis (LDA) can be used to reduce the dimensionality of the original data, thereby reducing data complexity. It can also utilize type label information to maximize the ratio of between-class variance to within-class variance, thus enhancing the separability of the data.
[0033] S2: Based on the feature dataset, the optimal parameters of the feature subset and the support vector machine are obtained through a multi-objective particle swarm optimization algorithm, specifically: S21: Initialize the parameters of the multi-objective particle swarm optimization algorithm. The parameters include the position, velocity, mutation rate exponent, and preset number of iterations of the particles. The position of each particle in the particle swarm is a combination of the parameters of the support vector machine and the features selected from the feature dataset. Different particle positions have different combinations of parameters and different combinations of features.
[0034] The search objectives for particles are twofold: to maximize classification accuracy and to minimize the number of features.
[0035] In addition, the parameters also include: inertia weight. Inertia used to control particle movement; First cognitive coefficient , representing the weight of the particle's own experience; the second cognitive coefficient , representing the weight of the particle swarm experience; the number of particles, and the upper and lower boundaries of the search space.
[0036] The parameters of a support vector machine, namely the penalty parameter C and the kernel function parameter γ, determine its classification performance.
[0037] The selection of features from a feature dataset can be achieved using a feature selection mask, which is a binary vector representing the selected feature.
[0038] In addition, a repository for non-dominated solutions needs to be created and the size of the repository needs to be initialized.
[0039] S22: Construct a fitness value vector based on classification accuracy and feature scores. Calculate the particle's fitness value vector based on the particle's position, specifically: According to the formula: Calculate the fitness value vector of the particles; In the formula, Let be the particle's fitness value vector. For classification accuracy, These are characteristic scores.
[0040] The classification accuracy is based on the formula: Calculate the classification accuracy; In the formula, The number of samples that are true negatives and correctly predicted as negative. The number of samples that are true positives and correctly predicted as positive. The number of positive samples that were incorrectly predicted as negative due to false negatives; The number of negative samples that were incorrectly predicted as positive (false positives). The feature score is calculated according to the formula: Calculate the characteristic scores; In the formula, The number of features selected from the feature dataset. This represents the total number of features in the feature dataset.
[0041] S23: If the particle's fitness value vector is better than the particle's historical best fitness value vector, then update the particle's best position according to the particle's fitness value vector. If a particle's fitness value vector is not better than its historical best fitness value vector, then the particle's best position is not updated.
[0042] In multi-objective particle swarm optimization (PSO), each particle has its own velocity and position. Velocity represents the particle's direction and speed of movement in the search space, while position represents the particle's current location in the current round. Each particle searches for the optimal solution independently in the search space and records it as its current individual extreme value, i.e., the particle's historical best fitness value vector, thus ensuring that each particle records the best solution encountered during its own search process.
[0043] S24: Determine the non-dominated solution based on the particle's optimal position; If the repository is not full, add non-dominated solutions directly to the repository; If the repository is full, calculate the crowding degree of non-dominated solutions in the repository, remove the non-dominated solution with the highest crowding degree, and add the new non-dominated solution to the repository.
[0044] Specifically, after obtaining the optimal position of a particle, it is determined whether the current particle's fitness value vector is better than some non-dominated solutions in the repository. If the particle's fitness value vector is better than some solutions in the repository in both the objectives of "maximizing classification accuracy and minimizing the number of features", then the particle needs to be added to the repository.
[0045] If the repository is full, the crowding degree of all non-dominated solutions in the repository is calculated. A roulette wheel selection strategy can be used to remove the non-dominated solution with the highest crowding degree to make room for new non-dominated solutions. A non-dominated solution with high crowding degree indicates that the solution has a high distribution density in the target space, which means it contributes less diversity to the repository. Removing non-dominated solutions with high crowding degree ensures that the repository always stores the best non-dominated solution in the current particle swarm and maintains diversity.
[0046] S25: The position of the particle with the least crowded non-dominated solution in the repository is taken as the optimal position of the particle swarm.
[0047] The optimal position of the particle swarm represents the highest fitness value vector of the particle, thereby maximizing classification accuracy and minimizing the number of features.
[0048] S26: Based on the particle's velocity, optimal particle position, and optimal particle swarm position in the current round, update the particle's velocity to obtain the updated particle velocity, specifically: According to the formula: Update the particle velocity to obtain the updated particle velocity; In the formula, k is the iteration round; Let be the velocity of the i-th particle in the (k+1)-th round; Let be the velocity of the i-th particle in the k-th round; Let be the optimal position of the i-th particle; Let i be the position of the i-th particle in the k-th round; Inertial weights are used to control the inertia of particle movement. The first cognitive coefficient represents the weight of the particle's own experience; The second cognitive coefficient represents the weight of the particle swarm's experience; and All are random numbers between 0 and 1; This is the optimal position for the particle swarm.
[0049] Then, based on the updated particle velocity and the particle position in the current round, the particle position is updated to obtain the particle's first position, specifically: According to the formula: Update the particle's position; In the formula, Let i be the position of the i-th particle in the (k+1)th round; Let i be the position of the i-th particle in the k-th round; Let be the velocity of the i-th particle in the (k+1)-th round.
[0050] S27: Based on the particle's first position, mutation rate exponent, preset iteration number, and current round, obtain the mutation probability, and generate a random number based on the mutation probability; If the random number is less than the mutation probability, then the first position of the particle is mutated to obtain the updated position of the particle. If the random number is not less than the mutation probability, then no mutation operation is performed, and the first position of the particle is used as the position of the updated particle.
[0051] Specifically, in step S27, the mutation probability is obtained based on the particle's first position, mutation rate exponent, preset iteration number, and current round, according to the formula: , thus obtaining the mutation probability; In the formula, The mutation probability, It is the variability index. To preset the number of iterations, This refers to the current round.
[0052] As the number of iterations increases, the value of the current round becomes larger, and the mutation probability decreases. That is, the mutation probability is higher in the early stages of the search and lower in the later stages. This helps to increase diversity in the early stages of the search, while the later stages focus more on local searches.
[0053] In each iteration, based on the current mutation probability Generate a random number rand. If rand < If the selected particle's position is positive, a mutation operation is performed; otherwise, no mutation is performed. For the selected particle, a dimension (i.e., an element in the position vector) is randomly chosen for mutation. The mutated particle's position can be randomly selected between upper and lower bounds, thereby changing the positions of some particles, increasing the diversity of solutions, and obtaining updated particle positions. The specific formula is as follows: ,in, Let i be the position of particle i. As the lower bound of the variable, The upper bound of the variable. This is the upper bound of the search space. This serves as the lower bound of the search space. The particle position... Mutate to a new location ,in yes Random values within the interval.
[0054] S28: If the preset conditions are met, the optimal parameters of the feature subset and the support vector machine are obtained based on the optimal position of the particle swarm. If the preset conditions are not met, execute S22 until the preset conditions are met, and obtain the optimal parameters of the feature subset and support vector machine based on the optimal position of the particle swarm.
[0055] Specifically, the preset condition can be a preset number of iterations or a preset fitness threshold.
[0056] S3: Match the optimal parameters to the support vector machine, construct a binary tree support vector machine based on binary tree theory, and obtain the anomaly detection model of the industrial control system after training with feature subsets.
[0057] By constructing a binary tree support vector machine, the multi-classification problem is transformed into multiple binary classification problems. Specifically, a binary tree structure is constructed, with each node being a binary support vector machine. Each binary support vector machine is trained on a feature subset using optimal parameters (C, γ). The binary tree support vector machines are used to detect and classify the test data to obtain prediction results. The prediction results are compared with the actual labels. The accuracy, recall, F1 score, and other indicators are calculated using a multi-class confusion matrix to evaluate the performance of the trained industrial control system anomaly detection model.
[0058] S4: Acquire the operating data of the industrial control system to be tested, and after the industrial control system anomaly detection model performs the test, output the type detection result.
[0059] Generally, type detection results include: Simple Malicious Response Injection (NMRI): This involves relatively simple malicious response injection; Complex Malicious Response Injection (CMRI): This involves complex malicious response injection, which may involve multiple steps or more complex logic. Malicious State Command Injection (MSCI): Changes the state of the system by injecting malicious commands; Malicious Parameter Injection (MPCI): This technique injects malicious parameters to affect the system's command execution. Malicious Functional Command Injection (MFCI): Injects malicious functional commands that may alter or disrupt system functions; Denial-of-Service (DoS) attack: An attack that renders a system unable to provide normal service by sending a large number of requests or other means; Recon attack: An attacker attempts to gather information about the system in order to launch a further attack.
[0060] The industrial control system anomaly detection method provided by this invention first standardizes and reduces the dimensionality of features to reduce the computer's workload. Then, it uses a particle swarm optimization algorithm to optimize the parameters of the support vector machine and select features, thereby obtaining both the optimal parameters of the support vector machine and a feature subset related to the anomaly type. The industrial control system anomaly detection model obtained based on the optimal parameters of the support vector machine and the feature subset has high operating efficiency and high accuracy in type prediction, and can accurately classify different types of attacks in the industrial control system.
[0061] The present invention also provides an industrial control system anomaly type detection system, comprising: Feature dataset construction module: used to acquire industrial control system operation data, including: tags, command addresses, response addresses, command memory counts, communication read functions, response write functions, sub-functions, setpoints, control modes, control schemes, and measured values. It calculates the standard deviation of features in the industrial control system operation data, deletes features with a standard deviation of zero, scales the features to have a mean of zero and a variance of 1, performs dimensionality reduction processing, and constructs the feature dataset. Optimization module: Based on the feature dataset, the optimal parameters of the feature subset and support vector machine are obtained through a multi-objective particle swarm optimization algorithm, specifically: The parameters of the multi-objective particle swarm optimization algorithm are initialized. The parameters include the position, velocity, mutation rate exponent, and preset number of iterations of the particles. The position of each particle in the particle swarm is a combination of the parameters of the support vector machine and the features selected from the feature dataset. Different particle positions have different parameter combinations and different feature combinations. A fitness value vector is constructed based on classification accuracy and feature scores, and the fitness value vector of a particle is calculated based on the particle's position. If a particle's fitness value vector is better than its historical best fitness value vector, then the particle's optimal position is updated based on the particle's fitness value vector. If a particle's fitness value vector is not better than its historical best fitness value vector, then the particle's best position is not updated. Determine the non-dominated solution based on the particle's optimal position; If the repository is not full, add non-dominated solutions directly to the repository; If the repository is full, calculate the crowding degree of non-dominated solutions in the repository, remove the non-dominated solution with the highest crowding degree, and add the new non-dominated solution to the repository. The position of the particle with the least crowded non-dominated solution in the repository is taken as the optimal position of the particle swarm. Based on the particle's velocity, optimal particle position, and optimal particle swarm position in the current round, update the particle's velocity to obtain the updated particle velocity. Update the particle's position based on the updated particle's velocity and the particle's position in the current round to obtain the particle's first position; Based on the particle's first position, mutation rate exponent, preset iteration number, and current round, the mutation probability is obtained, and a random number is generated based on the mutation probability. If the random number is less than the mutation probability, then the first position of the particle is mutated to obtain the updated position of the particle. If the random number is not less than the mutation probability, then no mutation operation is performed, and the first position of the particle is used as the position of the updated particle. If the preset conditions are met, the optimal parameters of the feature subset and the support vector machine are obtained based on the optimal position of the particle swarm. If the preset conditions are not met, the fitness value vector of the particles is recalculated until the preset conditions are met. Based on the optimal position of the particle swarm, the optimal parameters of the feature subset and the support vector machine are obtained. The industrial control system anomaly detection model establishment module is used to match the optimal parameters to the support vector machine. Based on the binary tree theory, a binary tree support vector machine is constructed. After training with feature subsets, the industrial control system anomaly detection model is obtained. Type detection module: Used to acquire the operating data of the industrial control system to be detected, and after being detected by the industrial control system anomaly detection model, output the type detection result.
[0062] The present invention also provides an industrial control system anomaly detection device, including a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the industrial control system anomaly detection method.
[0063] The present invention also provides an industrial control system anomaly detection medium for storing a computer program, wherein the computer program, when executed by a processor, implements the industrial control system anomaly detection method.
Claims
1. A method for detecting anomaly types in an industrial control system, characterized in that, Includes the following steps: S1: Acquire industrial control system operation data, including: tags, command addresses, response addresses, command memory counts, communication read functions, response write functions, sub-functions, setpoints, control modes, control schemes, and measured values. Calculate the standard deviation of features in the industrial control system operation data, delete features with a standard deviation of zero, scale the features to have a mean of zero and a variance of 1, perform dimensionality reduction processing, and construct a feature dataset. S2: Based on the feature dataset, the optimal parameters of the feature subset and the support vector machine are obtained through a multi-objective particle swarm optimization algorithm, specifically: S21: Initialize the parameters of the multi-objective particle swarm optimization algorithm. The parameters include the position, velocity, mutation rate exponent, and preset number of iterations of the particles. The position of each particle in the particle swarm is a combination of the parameters of the support vector machine and the features selected from the feature dataset. Different particle positions have different combinations of parameters and different combinations of features. S22: Construct a fitness value vector based on classification accuracy and feature scores, and calculate the particle's fitness value vector based on the particle's position; S23: If the particle's fitness value vector is better than the particle's historical best fitness value vector, then update the particle's best position according to the particle's fitness value vector. If a particle's fitness value vector is not better than its historical best fitness value vector, then the particle's best position is not updated. S24: Determine the non-dominated solution based on the optimal position of the particle; If the repository is not full, add non-dominated solutions directly to the repository; If the repository is full, calculate the crowding degree of non-dominated solutions in the repository, remove the non-dominated solution with the highest crowding degree, and add the new non-dominated solution to the repository. S25: The position of the particle with the least crowded non-dominated solution in the repository is taken as the optimal position of the particle swarm; S26: Based on the particle's velocity, optimal particle position, and optimal particle swarm position in the current round, update the particle's velocity to obtain the updated particle velocity; Update the particle's position based on the updated particle's velocity and the particle's position in the current round to obtain the particle's first position; S27: Based on the particle's first position, mutation rate exponent, preset iteration number, and current round, obtain the mutation probability, and generate a random number based on the mutation probability; If the random number is less than the mutation probability, then the first position of the particle is mutated to obtain the updated position of the particle. If the random number is not less than the mutation probability, then no mutation operation is performed, and the first position of the particle is used as the position of the updated particle. S28: If the preset conditions are met, the optimal parameters of the feature subset and the support vector machine are obtained based on the optimal position of the particle swarm. If the preset conditions are not met, execute S22 until the preset conditions are met, and obtain the feature subset and the optimal parameters of the support vector machine based on the optimal position of the particle swarm. S3: Match the optimal parameters to the support vector machine, construct a binary tree support vector machine based on binary tree theory, and obtain the anomaly detection model of the industrial control system after training with feature subsets; S4: Acquire the operating data of the industrial control system to be tested, and after the industrial control system anomaly detection model performs the test, output the type detection result.
2. The method for detecting anomaly types in industrial control systems according to claim 1, characterized in that, In step S22, a fitness value vector is constructed based on classification accuracy and feature scores. The fitness value vector of a particle is calculated based on its position, specifically as follows: According to the formula: Calculate the fitness value vector of the particles; In the formula, Let be the particle's fitness value vector. For classification accuracy, These are characteristic scores.
3. The method for detecting anomaly types in industrial control systems according to claim 2, characterized in that, The classification accuracy is based on the formula: Calculate the classification accuracy; In the formula, The number of samples that are true negatives and correctly predicted as negative. The number of samples that are true positives and correctly predicted as positive. The number of positive samples that were incorrectly predicted as negative due to false negatives; The number of negative samples that were incorrectly predicted as positive (false positives). The feature score is calculated according to the formula: Calculate the characteristic scores; In the formula, The number of features selected from the feature dataset. This represents the total number of features in the feature dataset.
4. The method for detecting anomaly types in industrial control systems according to claim 1, characterized in that, In step S26, the particle velocity is updated based on the particle velocity, optimal particle position, and optimal particle swarm position in the current round, resulting in the updated particle velocity. Specifically: According to the formula: Update the particle velocity to obtain the updated particle velocity; In the formula, k is the iteration round; Let be the velocity of the i-th particle in the (k+1)-th round; Let be the velocity of the i-th particle in the k-th round; Let be the optimal position of the i-th particle; Let i be the position of the i-th particle in the k-th round; Inertial weights are used to control the inertia of particle movement. The first cognitive coefficient represents the weight of the particle's own experience; The second cognitive coefficient represents the weight of the particle swarm's experience; and All are random numbers between 0 and 1; This is the optimal position for the particle swarm.
5. The method for detecting anomaly types in industrial control systems according to claim 1, characterized in that, In step S26, the particle's position is updated based on the updated particle velocity and the particle's position in the current round to obtain the particle's first position, specifically as follows: According to the formula: Update the particle's position; In the formula, Let i be the position of the i-th particle in the (k+1)th round; Let i be the position of the i-th particle in the k-th round; Let be the velocity of the i-th particle in the (k+1)-th round.
6. The method for detecting anomaly types in an industrial control system according to claim 1, characterized in that, In step S27, the mutation probability is obtained based on the particle's first position, mutation rate exponent, preset iteration number, and current round, according to the formula: , thus obtaining the mutation probability; In the formula, The mutation probability, It is the variability index. To preset the number of iterations, This refers to the current round.
7. The method for detecting anomaly types in industrial control systems according to claim 1, characterized in that, In S1, the features are scaled to have a mean of zero and a variance of 1, specifically as follows: According to the formula: , which scales the features to have a mean of zero and a variance of 1; In the formula, These are the standardized eigenvalues. These are the original eigenvalues; Original eigenvalues The mean; Original eigenvalues The standard deviation.
8. An industrial control system anomaly type detection system, characterized in that, include: Feature dataset construction module: used to acquire industrial control system operation data, including: tags, command addresses, response addresses, command memory counts, communication read functions, response write functions, sub-functions, setpoints, control modes, control schemes, and measured values. It calculates the standard deviation of features in the industrial control system operation data, deletes features with a standard deviation of zero, scales the features to have a mean of zero and a variance of 1, performs dimensionality reduction processing, and constructs the feature dataset. Optimization module: Based on the feature dataset, the optimal parameters of the feature subset and support vector machine are obtained through a multi-objective particle swarm optimization algorithm, specifically: The parameters of the multi-objective particle swarm optimization algorithm are initialized. The parameters include the position, velocity, mutation rate exponent, and preset number of iterations of the particles. The position of each particle in the particle swarm is a combination of the parameters of the support vector machine and the features selected from the feature dataset. Different particle positions have different parameter combinations and different feature combinations. A fitness value vector is constructed based on classification accuracy and feature scores, and the fitness value vector of a particle is calculated based on the particle's position. If a particle's fitness value vector is better than its historical best fitness value vector, then the particle's best position is updated based on the particle's fitness value vector. If a particle's fitness value vector is not better than its historical best fitness value vector, then the particle's best position is not updated. Determine the non-dominated solution based on the particle's optimal position; If the repository is not full, add non-dominated solutions directly to the repository; If the repository is full, calculate the crowding degree of non-dominated solutions in the repository, remove the non-dominated solution with the highest crowding degree, and add the new non-dominated solution to the repository. The position of the particle with the least crowded non-dominated solution in the repository is taken as the optimal position of the particle swarm. Based on the particle's velocity, optimal particle position, and optimal particle swarm position in the current round, update the particle's velocity to obtain the updated particle velocity. Update the particle's position based on the updated particle's velocity and the particle's position in the current round to obtain the particle's first position; Based on the particle's first position, mutation rate exponent, preset iteration number, and current round, the mutation probability is obtained, and a random number is generated based on the mutation probability. If the random number is less than the mutation probability, then the first position of the particle is mutated to obtain the updated position of the particle. If the random number is not less than the mutation probability, then no mutation operation is performed, and the first position of the particle is used as the position of the updated particle. If the preset conditions are met, the optimal parameters of the feature subset and the support vector machine are obtained based on the optimal position of the particle swarm. If the preset conditions are not met, the fitness value vector of the particles is recalculated until the preset conditions are met. Based on the optimal position of the particle swarm, the optimal parameters of the feature subset and the support vector machine are obtained. The industrial control system anomaly detection model establishment module is used to match the optimal parameters to the support vector machine. Based on the binary tree theory, a binary tree support vector machine is constructed. After training with feature subsets, the industrial control system anomaly detection model is obtained. Type detection module: Used to acquire the operating data of the industrial control system to be detected, and after being detected by the industrial control system anomaly detection model, output the type detection result.
9. An industrial control system anomaly detection device, characterized in that, It includes a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the industrial control system anomaly type detection method as described in any one of claims 1-7.
10. A detection medium for anomaly types in an industrial control system, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the industrial control system anomaly type detection method as described in any one of claims 1-7.
Citation Information
Patent Citations
Methods and compositions for highly sensitive detection of molecules
CN102016552A
PSO-OCSVM based industrial control system communication behavior anomaly detection method
CN105703963A