Self-checking method, system and equipment for network security of industrial control system

The self-check method for ICS uses K-means clustering to establish behavior baselines and detect anomalies, enhancing security by reducing false alarms and ensuring system stability.

CN120321024APending Publication Date: 2025-07-15PIPECHINA SOUTH CHINA CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510674013.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing industrial control system network security measures are difficult to effectively monitor equipment behavior and identify internal abnormalities in the network. Especially when facing complex attack methods, traditional protection methods are not enough to meet the needs.

Method used

By periodically collecting communication data from the industrial control system, establishing a behavior baseline, extracting feature values using the K-mean clustering algorithm, calculating deviations and comparing them with preset thresholds, generating alarm reports, and real-time abnormality detection and monitoring are achieved.

Benefits of technology

Real-time monitoring of the communication behavior of industrial control system equipment is realized, and potential abnormalities are identified in a timely manner, the accuracy of abnormal detection and system stability are improved, and the risk of safety accidents is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321024A_ABST
    Figure CN120321024A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data communication security, and discloses a self-checking method, system and equipment for network security of an industrial control system, and the method comprises the steps: periodically collecting communication data of the industrial control system in a normal operation environment, and storing historical data, modeling the historical data based on K-means clustering and extracting characteristic values to form behavior baselines of all devices in the industrial control system; comparing the communication data with a behavior baseline, calculating the deviation degree of each type of data in the communication data relative to the behavior baseline, comparing and judging the deviation degree with a preset deviation threshold value, judging whether abnormity exists or not according to the communication data and equipment information, and if the abnormity exists, sending the communication data to the behavior baseline; if yes, generating an alarm report. The behavior characteristics of each device and communication flow in the industrial control system are monitored in real time, abnormal behaviors are detected, and an alarm is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data communication security, and in particular, to a self-checking method, system and device for industrial control system network security. Background Art

[0002] With the wide application of industrial control systems (ICS) in modern manufacturing and infrastructure, their network security issues have attracted increasing attention. Traditional industrial control systems mostly rely on dedicated hardware and closed network architectures. However, with the rapid development of informatization and automation technologies, more and more industrial control systems are connected to enterprise internal networks, the Internet, and the industrial Internet of Things (IIoT). The security of these systems faces unprecedented challenges. The network security of industrial control systems is not only related to the confidentiality, integrity, and availability of data, but also directly affects the safety and reliability of the production process. Any system failure or attack can lead to serious economic losses, production stagnation, and even safety accidents. Therefore, improving the network security of industrial control systems, especially the real-time monitoring and protection of communication data and device behaviors, has become an urgent task.

[0003] In the current security protection of industrial control systems, traditional network security means (such as firewalls, intrusion detection, etc.) often cannot effectively cope with emerging security threats. Existing security mechanisms mostly focus on the defense against external intrusions, but are insufficient in the real-time monitoring of device behaviors and the identification of internal network anomalies. With the increasing complexity of attack means, simple boundary protection can no longer meet the security detection requirements for the internal behaviors of industrial control systems. Therefore, how to model device communication behaviors through intelligent means, identify abnormal behaviors, and give real-time warnings has become a technical problem in the current field. Summary of the Invention

[0004] In view of this, the present invention proposes a self-checking method, system and device for industrial control system network security, aiming to solve the problem of insufficient real-time monitoring of device behaviors and identification of internal network anomalies in the current technology.

[0005] A self-checking method for industrial control system network security proposed by the present invention includes:

[0006] Periodically collect communication data in the normal operating environment of the industrial control system and store historical data. The communication data includes five types: packet size, communication frequency, data transmission direction, communication protocol, and device operation log;

[0007] After cleaning the historical data, model the historical data based on K-means clustering and extract feature values to form the behavior baselines of each device in the industrial control system;

[0008] Compare the communication data with the behavior baseline, calculate the deviation degree of each type of data in the communication data compared with the behavior baseline, compare and judge the deviation degree with a preset deviation threshold. When the deviation degree exceeds the deviation threshold, it is judged that the device is abnormal and the device information of the device is recorded. The device information includes: device ID, abnormal type, timestamp and deviation degree;

[0009] Judge whether there is an abnormality according to the communication data and device information. If there is an abnormality, generate an alarm report;

[0010] Among them, when modeling the historical data based on K-means clustering and extracting feature values to form the behavior baseline of each device in the industrial control system, it includes:

[0011] Perform standardization processing on each piece of historical data, convert it into a standard normal distribution to eliminate the scale difference between features;

[0012] Use the elbow method to determine the optimal K value of clustering, calculate the change of clustering error under different K values, and select the point where the error decline rate slows down as the K value;

[0013] Input the feature vector of the device into the K-means algorithm for iterative optimization until the change of the clustering center is less than the set threshold. Each device is assigned to a cluster according to its communication behavior characteristics, and the clustering center is the behavior baseline of the device;

[0014] The behavior baseline of each device is determined by the feature values of its affiliated clustering center.

[0015] Further, when performing data cleaning on the historical data, it includes:

[0016] Delete the records corresponding to the independently existing missing values and outliers;

[0017] If there are two or more consecutive missing values or outliers, record the corresponding timestamps respectively and judge to fill the missing values or correct the outliers.

[0018] Further, when performing data cleaning on the historical data, it also includes:

[0019] When judging whether to fill the missing values, it depends on whether there is a linear relationship between the historical data on both sides of the missing values; if there is, fill the missing values by linear interpolation method, if not, abandon filling the missing values;

[0020] When determining whether to correct the outlier, outlier detection is performed based on the statistical characteristics and trends of the historical data; when the outlier is less than or equal to the mean of the historical data + 3 times the standard deviation and greater than or equal to the mean of the historical data - 3 times the standard deviation, and there is a linear relationship between the historical data on both sides of the outlier, the outlier is corrected by linear interpolation method.

[0021] Further, before modeling the historical data based on K-means clustering and extracting eigenvalue to form the behavior baseline of each device in the industrial control system, it also includes:

[0022] Original feature extraction: Statistically analyze the packet size of each device, calculate the frequency of data sent / received by the device per unit time, record the upload / download traffic ratio of each device, analyze the protocol type and frequency used by each device, extract the frequency, type of device operations, and the duration of each operation;

[0023] Time series feature extraction: Use the sliding window method to perform time series modeling on the communication data of each device, and extract the change trend of the communication data;

[0024] Periodicity analysis: Use Fourier transform to extract the periodic characteristics of the packet size, communication frequency, data transmission direction, communication protocol, and device operation logs, and identify potential periodic patterns;

[0025] Outlier detection: Use the gradient-based algorithm to detect the outlier in the time series and capture the change pattern of abnormal behavior.

[0026] Further, when comparing and judging the deviation degree with a preset deviation threshold, it includes:

[0027] For each collected communication data Xi = [X1, X2, X3, X4, X5]; calculate the deviation degree between the communication data Xi and the corresponding behavior baseline Bi = [B1, B2, B3, B4, B5], and the deviation degree has the following relationship:

[0028]

[0029] Among them, Xi represents the original value of the i-th type of communication data feature; μ i represents the historical data mean of the i-th type of feature; σ i represents the historical data standard deviation of the i-th type of feature; d i represents the deviation degree of the i-th type of feature.

[0030] Further, compare the deviation degree with a preset deviation threshold. When the deviation degree exceeds the deviation threshold, it is determined that the device is abnormal and the device information of the device is recorded

[0031]

[0032] Compare the total deviation D with a preset deviation threshold T to determine whether there is an anomaly; if D > T, it is considered that the communication data of the device is abnormal;

[0033] where D is the total deviation.

[0034] Furthermore, it includes: when the communication data of the device is abnormal, based on the update mechanism of time series data and periodic characteristics, the behavior baseline of the device is updated again.

[0035] Furthermore, when there is an anomaly and an alarm report is generated, it includes:

[0036] When the total deviation D is less than the deviation threshold T, it is determined that the device is at a minor deviation level and a report is generated;

[0037] When the total deviation D is greater than or equal to the deviation threshold T, it is determined that the device is at a serious deviation level and a report is generated.

[0038] In some embodiments, the present invention also provides a self-checking system for industrial control system network security, which is applied to the above method. The system includes:

[0039] A data acquisition module, configured to periodically acquire the communication data of devices in the industrial control system, including the packet size, communication frequency, data transmission direction, communication protocol, and device operation logs, and store the historical data;

[0040] A data cleaning and preprocessing module, configured to clean the acquired historical data, delete missing values and outliers, and perform necessary data correction and filling to ensure the integrity and accuracy of the data;

[0041] A behavior baseline modeling module, configured to model the historical data based on the K-means clustering algorithm, extract the communication behavior characteristics of the devices, and form the behavior baseline of each device for subsequent anomaly detection;

[0042] An anomaly detection module, configured to compare the real-time acquired communication data with the behavior baseline of the device, calculate the deviation of the data, and determine whether there is a device anomaly according to the preset deviation threshold;

[0043] An alarm generation module, configured to generate an alarm report when a device anomaly is detected. The report includes device information, anomaly type, timestamp, and deviation, and send an alarm notification to the user.

[0044] Compared with the prior art, the beneficial effects of the present application are as follows:

[0045] By periodically collecting and analyzing the communication data of each device in the industrial control system and establishing a behavior baseline, the communication behavior of the device can be monitored in real time. Once a deviation of the communication data from the normal behavior baseline is detected, potential abnormal behaviors or security threats can be identified in a timely manner, ensuring the stable operation of the industrial control system and reducing the risks of cyber attacks, equipment failures, or malicious behaviors.

[0046] Adopting the K-means clustering algorithm for historical data modeling and behavior baseline extraction can accurately capture the normal communication patterns of devices based on large-scale historical data. Based on this baseline, the system can accurately calculate the deviation degree of each device and compare it with a preset threshold to accurately determine whether the device is abnormal, avoiding false negatives and false positives of traditional detection methods and improving the accuracy of anomaly detection.

[0047] During the data cleaning process, the system can automatically detect and correct missing values and outliers to ensure data quality. The missing values and outliers are filled and corrected by the linear interpolation method to ensure the continuity and consistency of the data, effectively improving the reliability of subsequent modeling and anomaly detection.

[0048] This solution also introduces an update mechanism based on time series data and periodic characteristics, which can automatically update the behavior baseline when the communication data of the device is abnormal. This mechanism enables the system to cope with possible changes in the industrial control system environment (such as new devices added, configuration changes, etc.), ensuring that the behavior baseline keeps pace with the times and improving the flexibility and adaptability of the system.

[0049] The system can not only generate alarm reports according to the deviation degree, but also distinguish the types of anomalies (slight deviation or severe deviation) according to the severity of the deviation, so as to more accurately guide the staff to take different response measures. Slight deviations can be observed first, while severe deviations will trigger an emergency response to ensure timely handling of potential problems.

[0050] The system can automatically collect communication data, analyze data, build a behavior baseline model and monitor the status of devices in real time, greatly reducing the pressure of manual monitoring. Alarm reports can be generated in a timely manner and sent to relevant personnel to ensure that problems can be quickly discovered and responded to at the initial stage of their occurrence.

[0051] By continuously monitoring and analyzing the communication data, potential anomalies and security problems can be discovered and repaired in a timely manner, thereby improving the stability and reliability of the industrial control system. In addition, the anomaly detection and alarm generation mechanism makes the maintenance of the system more proactive and intelligent, helping to prevent problems in advance rather than dealing with faults afterwards.

[0052] In some embodiments, the present application provides an electronic device, including: a processor and a memory configured to store instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement any of the above optional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0054] Figure 1 It is a flowchart of a self-checking method for industrial control system network security according to an embodiment of the present invention;

[0055] Figure 2 It is a functional framework diagram of a self-checking system for industrial control system network security according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0057] In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "front", "rear", "inner", "outer", etc. is based on the orientation or relative positional relationship shown in the drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present application. Without special instructions, in the case of satisfying the relative positional relationship shown in the drawings, the above directional descriptions can be flexibly set during the actual application process.

[0058] The terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "plurality" is two or more.

[0059] In the description of this application, it should be noted that unless otherwise clearly specified and defined, the terms "installation", "connection", "linkage", and "communication" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection. It can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0060] In the embodiments of this application, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, article, or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the existence of additional identical elements in the process, article, or device including that element.

[0061] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0062] In the description of this specification, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.

[0063] Refer to Figure 1 As shown, the embodiments of the present invention provide a self-checking method for industrial control system network security, including:

[0064] S1: Periodically collect the communication data in the normal operating environment of the industrial control system and store the historical data. The communication data includes five types: packet size, communication frequency, data transmission direction, communication protocol, and device operation log;

[0065] S2: After cleaning the historical data, model the historical data based on K-means clustering and extract the feature values to form the behavior baselines of each device in the industrial control system;

[0066] S3: Compare the communication data with the behavior baselines, calculate the deviation degrees of various types of data in the communication data compared with the behavior baselines, compare the deviation degrees with a preset deviation threshold for judgment. When the deviation degree exceeds the deviation threshold, it is judged that the device is abnormal and record the device information of the device. The device information includes: device ID, abnormal type, timestamp, and deviation degree;

[0067] S4: Determine whether there is an anomaly based on the communication data and device information. If there is an anomaly, generate an alarm report.

[0068] It can be understood that by periodically collecting and storing the communication data in the normal operating environment of the industrial control system, the communication behavior patterns of the devices and the system can be comprehensively understood. These data provide a basis for subsequent behavior baseline modeling, ensuring that any anomalies inconsistent with normal behavior can be effectively detected, thereby improving the security and reliability of the system and timely discovering potential security threats.

[0069] Using the K-means clustering method to model historical data can extract the communication behavior characteristics of each device and form a behavior baseline. By comparing with the real-time collected communication data, calculating the deviation degree and comparing it with the preset deviation threshold, it can effectively determine whether there is an anomaly in the device. This anomaly detection method based on the deviation degree can quickly discover potential problems in the device operation, avoid delaying the detection of faults, take response measures in advance, and thus reduce security accidents caused by device anomalies.

[0070] When the deviation degree exceeds the preset threshold, the system can accurately identify the abnormal device and record the detailed information of the device (such as device ID, anomaly type, timestamp, and deviation degree). This process not only helps to quickly locate the problem device but also provides necessary basis for subsequent fault troubleshooting and security auditing, improving the response efficiency and problem-solving speed.

[0071] By cleaning and correcting historical data, eliminating the interference of missing values and outliers on the modeling process, the accuracy and integrity of the data are ensured. This data cleaning and correction mechanism can guarantee the high quality of the modeling results, make the construction of the behavior baseline more accurate, and thus enhance the effect of anomaly detection.

[0072] Based on the combination of the K-means clustering algorithm and deviation degree judgment, it helps to accurately distinguish normal fluctuations and real abnormal behaviors. This can effectively avoid false alarms and missed alarms, reduce unnecessary alarms, reduce the workload of personnel, and improve the efficiency and credibility of the alarm system.

[0073] When the communication data of the device is abnormal, the system automatically adjusts and updates the behavior baseline of the device through the update mechanism of time series data and periodic characteristics. This enables the system to adaptively optimize the behavior baseline, continuously improve the accuracy of anomaly detection as the device operating environment changes, and ensure the stable operation of the system over a long period of time.

[0074] In some embodiments of the present application, when cleaning the historical data, it includes:

[0075] Delete the records corresponding to the independently existing missing values and outliers;

[0076] If there are two or more consecutive missing values or outliers, record the corresponding timestamps respectively and determine whether to fill in the missing values or correct the outliers.

[0077] It can be understood that during data cleaning, it is necessary to first identify the missing values and outliers in the data. Missing values usually refer to the situation where some fields in the data records have no valid values, while outliers refer to those values that significantly deviate from the normal data range. Deleting these missing values and outliers can prevent these abnormal data from affecting the subsequent data analysis and modeling processes, but it is necessary to ensure that the amount of data deleted is not excessive, otherwise it may affect the representativeness of the data.

[0078] When missing values or outliers appear continuously, simple deletion may lead to data gaps. At this time, recording the timestamps can trace back to the specific time points, which is convenient for subsequent processing. When judging whether to fill in the missing values or correct the outliers, it can be decided according to the overall trend of the data. If the data change trends on both sides are consistent, interpolation can be used for filling. If the trends are inconsistent, filling or correction may need to be abandoned.

[0079] For consecutive missing values, the relationship between the front and back data should be considered when filling. If the front and back data show a linear relationship, linear interpolation can be used to fill in the missing values, so as to maintain the continuity and consistency of the data.

[0080] For outliers, statistical methods (such as mean ± 3 times standard deviation) can be used to determine the abnormal range. If the outliers do not meet these criteria but are consistent with the data change trends on both sides, these outliers can be corrected by interpolation to prevent them from affecting the subsequent modeling and analysis.

[0081] In some embodiments of the present application, when cleaning historical data, it further includes:

[0082] When judging whether to fill in the missing values, according to whether there is a linear relationship between the historical data on both sides of the missing values; if there is, fill in the missing values by linear interpolation, if not, abandon filling in the missing values;

[0083] When judging whether to correct the outliers, perform outlier detection based on the statistical characteristics and trends of the historical data; when the outlier is less than or equal to the mean of the historical data + 3 times the standard deviation and greater than or equal to the mean of the historical data - 3 times the standard deviation, and there is a linear relationship between the historical data on both sides of the outlier, correct the outlier by linear interpolation.

[0084] It should be noted that when judging whether to fill in missing values: when there are missing values in the data, it is first necessary to determine whether there is a linear relationship between the historical data on both sides of the missing value. This is because in many cases, data changes may show a certain trend or pattern. If the data on both sides show a consistent trend, it is reasonable to use linear interpolation to fill in the missing values, as it can maintain the smoothness and continuity of the data.

[0085] Linear interpolation method: Linear interpolation is a method of estimating missing data through the linear relationship of known data points. If there is a linear relationship between the data on both sides, the interpolation method will calculate the missing value through the straight-line formula to ensure that the change trend of the data is consistent with the historical data.

[0086] Give up filling: If the historical data on both sides of the missing value do not have a linear relationship (for example, the data fluctuates greatly and no clear trend can be formed), it is not recommended to use the interpolation method to fill in the missing value. This is because the filled data may cause errors in subsequent analysis and even affect the accuracy of the model. In this case, it is usually chosen not to fill in the missing value or to exclude it during data analysis.

[0087] Judgment based on statistical characteristics: The correction of outliers is detected based on the statistical characteristics (such as mean, standard deviation, etc.) and trends of historical data. Outliers usually refer to those data points that significantly deviate from the normal data range. By calculating the range of mean ± 3 times the standard deviation, it is possible to initially determine which data points belong to outliers.

[0088] Correction conditions: For the data points determined to be outliers, only when its value is less than or equal to the mean of historical data + 3 times the standard deviation and greater than or equal to the mean - 3 times the standard deviation, and there is a linear relationship between the data on both sides of the outlier, will correction be considered. This correction condition combines the statistical standard (mean ± 3 times the standard deviation) with the data trend (linear relationship) to ensure that corrections do not misprocess extreme outliers.

[0089] Correct outliers using the linear interpolation method: When the outliers meet the above conditions, the linear interpolation method is used to correct them. Through the interpolation method, the outliers will be adjusted to reasonable values that conform to the data trends before and after them. This can avoid the negative impact of these outliers on subsequent behavior baseline modeling or anomaly detection processes.

[0090] It can be understood that during the data cleaning process, missing values and outliers are effectively processed, avoiding the negative impact of missing data or incorrect data on subsequent analysis and modeling. Especially by using the linear interpolation method to fill in missing values and correct outliers, the smoothness and consistency of the data can be ensured, thereby improving the integrity and accuracy of the data.

[0091] For missing values that cannot be filled by interpolation, the "abandon filling" strategy is adopted, which can avoid inappropriate correction of data, prevent errors caused by artificial filling, and thus maintain the authenticity and reliability of data.

[0092] The historical data after data cleaning conforms more to statistical laws and avoids data noise caused by missing values or outliers. This is crucial for subsequent behavior baseline modeling and anomaly detection, and can ensure that algorithms such as K-means clustering can more accurately capture the normal behavior patterns of devices when performing device behavior modeling.

[0093] By correcting outliers and filling missing values, the trends and fluctuations of the data are better maintained, enabling the behavior baseline to truly reflect the communication behavior of the device, thereby improving the accuracy and reliability of anomaly detection.

[0094] The judgment criteria for outliers and missing values (such as mean ± 3 times the standard deviation and the judgment of linear relationships) can effectively reduce misjudgment and missed judgment. Only when outliers or missing values meet specific conditions are they processed, which ensures the accuracy of data correction operations and avoids interference with irrelevant data.

[0095] This correction method can ensure that the system will not be affected by incorrect corrected data when judging abnormal behaviors, improving the sensitivity and accuracy of anomaly detection.

[0096] By effectively cleaning and correcting historical data, it is possible to handle communication data anomalies caused by equipment failures, network problems, or other abnormal factors. Whether it is short-term data loss or occasional abnormal data fluctuations, the system can ensure data quality through a self-correction mechanism, thereby enhancing the robustness and fault tolerance of the entire system.

[0097] In a complex industrial environment, various problems in the data acquisition process may not be completely avoided. A good data cleaning and correction mechanism can help the system better adapt to and cope with these challenges.

[0098] Through strict preprocessing and correction of the data, the influence of extreme outliers on the data distribution is reduced, making subsequent model training based on these data (such as K-means clustering, deviation calculation, etc.) more stable. This can ensure the reliability of the device behavior baseline and enhance the stability of the entire system during long-term operation.

[0099] In some embodiments of the present application, before modeling historical data based on K-means clustering and extracting eigenvalue to form the behavior baseline of each device in the industrial control system, it further includes:

[0100] Original feature extraction: Statistically analyze the packet size of each device, calculate the frequency of data sent / received by the device per unit time, record the upload / download traffic ratio of each device, analyze the type and frequency of protocols used by each device, and extract the frequency, type, and duration of each device operation;

[0101] Time series feature extraction: Use the sliding window method to perform time series modeling on the communication data of each device and extract the trend of communication data changes;

[0102] Periodicity analysis: Use Fourier transform to extract the periodic characteristics of packet size, communication frequency, data transmission direction, communication protocol, and device operation logs, and identify potential periodic patterns;

[0103] Outlier detection: Use a gradient-based algorithm to detect outliers in the time series and capture the change patterns of abnormal behaviors.

[0104] It can be understood that before modeling historical data based on K-means clustering and extracting feature values to form the behavior baselines of each device in the industrial control system, a series of preprocessing steps were implemented, including original feature extraction, time series feature extraction, periodicity analysis, and outlier detection. The specific details are as follows:

[0105] Packet size: By statistically analyzing the packet size sent or received by each device, the communication load of the device and the network bandwidth occupancy can be understood.

[0106] Data transmission frequency per unit time: Calculating the frequency of data sent and received by the device per unit time helps monitor the activity level and normal operation status of the device.

[0107] Upload / download traffic ratio: Recording the ratio of device upload and download traffic can be used to analyze the communication mode of the device and whether there is abnormal data flow (for example, an abnormal large amount of upload may be a sign that the device is under attack).

[0108] Protocol type and frequency: Analyzing the communication protocols used by the device and their frequencies helps identify whether the device is using unauthorized or abnormal protocols, and this information is crucial for identifying potential security threats.

[0109] Frequency, type, and duration of device operations: Extracting the frequency, type, and duration of the operation behaviors of the device (such as startup, stop, and fault operations) can help understand whether the device is operating as expected and thus identify abnormal operations.

[0110] The sliding window method is used to model the communication data of each device, so as to extract the changing trend of the communication data. The sliding window method can help the system identify the long-term behavior changes of the device, such as the changes in the communication patterns of the device in different time periods, and then judge whether there are sudden anomalies or irregular behaviors.

[0111] The Fourier transform is used to extract the periodic characteristics of communication data (such as packet size, communication frequency, data transmission direction, etc.). This can help identify periodic patterns in device communication, such as timed communication behaviors, and then capture potential periodic attacks (such as sending abnormal requests at regular intervals), or discover regular changes in the normal operation of the device.

[0112] An algorithm based on gradient is used for outlier detection, and the gradient of the data is calculated to detect abnormal change patterns in the time series. The parts with significant gradient changes usually indicate that the device has abnormal behaviors at a certain moment or in a short period of time, such as a sharp increase in communication volume or abnormal data fluctuations. This method can effectively capture sudden or abnormal data that deviates from the normal pattern.

[0113] Through raw feature extraction, time series feature extraction and periodic analysis, the system can comprehensively understand the communication behaviors of each device, including the device's load, communication patterns, operation characteristics and their periodic changes. These features provide more accurate and rich input data for subsequent K-means clustering modeling, which helps to accurately establish the behavior baseline of the device.

[0114] Protocol type analysis and traffic ratio analysis can help identify unusual communication behaviors, such as unauthorized protocols or abnormal traffic patterns, which helps to detect potential security threats, such as malware, data leakage, DDoS attacks, etc.

[0115] Periodic analysis can help identify abnormal periodic behaviors in device operation by identifying potential periodic patterns in device communication, such as timed abnormal data transmission, periodic vulnerability exploitation, etc.

[0116] The outlier detection algorithm based on gradient can accurately capture the abnormal fluctuations in the device communication data, especially sudden abnormal changes. Through this method, the system can quickly identify the abnormal fluctuations in device operation, provide early warnings, and prevent potential security accidents from occurring.

[0117] Using the sliding window method to model the time series helps to extract the long-term changing trend of the device communication data, providing the system with the ability to predict the behavior of the device. This trend analysis can not only identify past anomalies, but also predict possible future abnormal behaviors of the device, so as to make responses in advance.

[0118] The data features extracted and analyzed through these preprocessing steps can ensure that the behavior baseline of each device is more in line with its actual operating conditions, reducing model bias. The accuracy of this baseline lays a solid foundation for subsequent deviation calculation and anomaly detection, improving the accuracy and reliability of anomaly detection.

[0119] Through comprehensive feature extraction and analysis, the normal behavior patterns of devices can be captured more accurately, avoiding misjudgment or missed judgment caused by a small amount of data or only considering a single feature, and improving the stability and credibility of the system.

[0120] In some embodiments of the present application, when modeling historical data based on K-means clustering and extracting feature values to form the behavior baseline of each device in the industrial control system, it includes:

[0121] Perform standardization processing on each historical data, converting it into a standard normal distribution to eliminate the scale difference between features;

[0122] Use the elbow method to determine the optimal K value for clustering, calculate the change in clustering error under different K values, and select the point where the error decline slows down as the K value;

[0123] Input the feature vector of the device into the K-means algorithm and perform iterative optimization until the change in the clustering center is less than the set threshold. Each device is assigned to a cluster according to its communication behavior characteristics, and the clustering center is the behavior baseline of the device;

[0124] The behavior baseline of each device is determined by the feature values of the cluster center it belongs to.

[0125] It can be understood that performing standardization processing on each historical data means converting the data into a standard normal distribution. The purpose of standardization processing is to eliminate the possible scale differences between different features. For example, the dimension and range differences between the packet size and communication frequency may cause the model to be biased towards a certain feature. Through standardization, the mean of each feature is 0 and the standard deviation is 1, making the influence of each feature on the clustering algorithm equal, thus avoiding errors caused by scale inconsistency.

[0126] The elbow method is a commonly used clustering analysis method for determining the number of clusters (K value) in the K-means algorithm. By calculating the clustering error (usually the sum of squared errors, SSE) under different K values and observing the trend of the error changing with the K value. When the K value increases, the clustering error gradually decreases, but after a certain point, the rate of error reduction slows down. This inflection point of the change is the elbow, and usually this point is selected as the optimal K value for clustering. Selecting the appropriate K value can ensure the best clustering effect, not only avoiding overfitting but also avoiding information loss caused by too few clusters.

[0127] Input the feature vectors of the devices into the K-means algorithm and perform iterative optimization. The K-means algorithm randomly initializes K cluster centers and assigns each device to the nearest cluster center according to its feature vector. Then, the algorithm recalculates the centers of each cluster until the change in the cluster centers is less than a set threshold, indicating that the clustering has converged and the eigenvalues no longer change. This process ensures that each device can be assigned to a suitable cluster according to its communication behavior characteristics.

[0128] The behavior baseline of each device is determined by the eigenvalue of the cluster center it belongs to. The cluster center represents the typical characteristics of the devices in this class in terms of communication behavior, such as packet size, communication frequency, traffic ratio, etc. The similarity between the device and the cluster center determines the "normal" range of its behavior. When the communication behavior of the device deviates from the cluster center, it can be determined as an abnormal behavior.

[0129] Normalization processing ensures that the scales of different features are consistent, so that the clustering algorithm will not be affected by some feature values with larger or smaller ranges, thereby improving the stability and accuracy of clustering. In this way, the system can more accurately identify the normal behavior patterns of devices and avoid errors caused by scale differences.

[0130] Using the elbow method can effectively determine the optimal value of K, avoiding overfitting or information loss caused by too many or too few clusters. Selecting an appropriate number of clusters helps to improve the accuracy of the behavior baseline, enabling the behavior baseline of each device to truly reflect its communication characteristics, and thus improving the accuracy of anomaly detection.

[0131] Through the iterative optimization process, the K-means algorithm continuously adjusts the cluster centers, enabling the communication behaviors of devices to be classified as accurately as possible. The convergence of this process ensures that each device is accurately divided into a cluster representing its behavior characteristics, improving the accuracy of the device behavior baseline.

[0132] The behavior baseline of each device is determined by the eigenvalue of the center of the cluster it belongs to. Therefore, when the communication behavior of the device deviates, the abnormal behavior can be identified by calculating the deviation degree between the device and the cluster center. The accuracy and reliability of the behavior baseline directly affect the precision of anomaly detection, ensuring the effectiveness of the industrial control system in security monitoring.

[0133] As an unsupervised learning method, K-means clustering does not rely on pre-labeled data and can flexibly adapt to industrial control systems of different types and scales. The system can automatically adjust its behavior baseline as new devices are added or existing devices change, and has good scalability.

[0134] In some embodiments of the present application, when comparing and judging the deviation degree with a preset deviation threshold, it includes:

[0135] For each collected communication data Xi = [X1, X2, X3, X4, X5]; calculate the deviation degree between the communication data Xi and the corresponding behavior baseline Bi = [B1, B2, B3, B4, B5], and the deviation degree has the following relationship:

[0136]

[0137] where Xi represents the original value of the communication data characteristics of the i-th type; μ i represents the historical data mean of the i-th type of characteristics; σ i represents the historical data standard deviation of the i-th type of characteristics; d i represents the deviation degree of the i-th type of characteristics.

[0138] It should be noted that by calculating the deviation degree, the system can carefully evaluate the difference between the communication behavior of each device and its historical behavior baseline. The deviation degree calculation method after standardization processing can effectively eliminate the dimension difference of different data characteristics, thus ensuring the accuracy of anomaly detection.

[0139] The calculation of the deviation degree helps the system quickly identify abnormal situations with large deviations from normal behavior in the communication data and generate alarm reports in a timely manner. This mechanism can ensure the sensitivity of the industrial control system to device anomalies during real-time monitoring, discover potential problems in advance, and reduce the risk of device failures.

[0140] By comparing the deviation degree with a preset threshold, the system can make flexible judgments according to the normal behavior patterns of different devices. With the change of the number and types of devices in the industrial control system, the deviation degree calculation method has good adaptability and can automatically adjust and identify the behavior patterns of new devices.

[0141] The calculation of the deviation degree does not depend on manually set rules, but makes judgments based on a data-driven model. This method can continuously optimize and improve the accuracy of anomaly detection and can dynamically adjust the detection criteria according to the historical data and behavior patterns of device operation.

[0142] Accurately detecting anomalies in device communication behavior can effectively prevent potential cyber attacks or device failures, and improve the security and stability of the industrial control system. By judging the deviation degree, it can help discover network security threats or device problems in a timely manner, thereby reducing the impact on the production process.

[0143] In some embodiments of the present application, the deviation degree is compared with a preset deviation threshold for judgment. When the deviation degree exceeds the deviation threshold, it is judged that the device is abnormal and the device information of the device is recorded.

[0144]

[0145] Compare the total deviation degree D with a preset deviation threshold T to determine whether there is an anomaly; if D > T, it is considered that the communication data of the device is abnormal;

[0146] Among them, D is the total deviation degree.

[0147] In some embodiments of the present application, it includes: when the communication data of the device is abnormal, based on the update mechanism of time series data and periodic characteristics, the behavior baseline of the device is updated again.

[0148] In some embodiments of the present application, when an anomaly exists and an alarm report is generated, it includes:

[0149] When the total deviation degree D is less than the deviation threshold T, it is determined that the device is in a minor deviation level and a report is generated;

[0150] When the total deviation degree D is greater than or equal to the deviation threshold T, it is determined that the device is in a serious deviation level and a report is generated.

[0151] Refer to Figure 2 As shown, the embodiment of the present invention provides a self-checking system for industrial control system network security, including:

[0152] A data acquisition module, configured to periodically collect the communication data of devices in the industrial control system, including the packet size, communication frequency, data transmission direction, communication protocol, and device operation logs, and store the historical data;

[0153] A data cleaning and preprocessing module, configured to clean the collected historical data, delete missing values and outliers, and perform necessary data correction and filling to ensure the integrity and accuracy of the data;

[0154] A behavior baseline modeling module, configured to model the historical data based on the K-means clustering algorithm, extract the communication behavior characteristics of the devices and form the behavior baseline of each device for subsequent anomaly detection;

[0155] An anomaly detection module, configured to compare the real-time collected communication data with the behavior baseline of the device, calculate the deviation degree of the data, and determine whether there is a device anomaly according to a preset deviation threshold;

[0156] An alarm generation module, configured to generate an alarm report when a device anomaly is detected. The report includes device information, anomaly type, timestamp, and deviation degree, and send an alarm notification to the user.

[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific implementation manners of the present invention, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A self-checking method for industrial control system network security, characterized in that, Including: Periodically collect communication data in the normal operating environment of the industrial control system and store historical data. The communication data includes five types: packet size, communication frequency, data transmission direction, communication protocol, and device operation log. After cleaning the historical data, model the historical data based on K-means clustering and extract eigenvalue to form the behavior baseline of each device in the industrial control system. Compare the communication data with the behavior baseline, calculate the deviation degree of each type of data in the communication data compared with the behavior baseline, compare the deviation degree with a preset deviation threshold for judgment. When the deviation degree exceeds the deviation threshold, it is judged that the device is abnormal and record the device information of this device. The device information includes: device ID, abnormal type, timestamp, and deviation degree. Judge whether there is an abnormality according to the communication data and device information. If there is an abnormality, generate an alarm report. Among them, when modeling the historical data based on K-means clustering and extracting eigenvalue to form the behavior baseline of each device in the industrial control system, it includes: Perform standardization processing on each historical data, convert it into a standard normal distribution to eliminate the scale difference between features. Use the elbow method to determine the optimal K value of clustering, calculate the change of clustering error under different K values, and select the point where the error decline rate slows down as the K value. Input the feature vector of the device into the K-means algorithm for iterative optimization until the change of the clustering center is less than the set threshold. Each device is assigned to a cluster according to its communication behavior characteristics, and the clustering center is the behavior baseline of this device. The behavior baseline of each device is determined by the eigenvalue of the clustering center it belongs to.

2. The method according to claim 1, wherein When cleaning the historical data, it includes: Delete the records corresponding to the independently existing missing values and outliers. If there are two or more consecutive missing values or outliers, record the corresponding timestamps respectively and judge to fill the missing values or correct the outliers.

3. The method according to claim 2, wherein When cleaning the historical data, it also includes: When judging whether to fill the missing values, according to whether there is a linear relationship between the historical data on both sides of the missing values. If there is, fill the missing values by linear interpolation method. If not, abandon filling the missing values. When judging whether to correct the outliers, perform outlier detection based on the statistical characteristics and trends of the historical data. When the outlier is less than or equal to the mean of the historical data + 3 times the standard deviation and greater than or equal to the mean of the historical data - 3 times the standard deviation, and there is a linear relationship between the historical data on both sides of the outlier, correct the outlier by linear interpolation method.

4. The method according to claim 3, characterized in that, Before modeling the historical data based on K-means clustering and extracting eigenvalue to form the behavior baseline of each device in the industrial control system, it also includes: Original feature extraction: Statistically analyze the packet size of each device, calculate the frequency of data sent / received by the device per unit time, record the upload / download traffic ratio of each device, analyze the protocol type and frequency used by each device, extract the frequency, type of device operations, and the duration of each operation. Time series feature extraction: Use the sliding window method to perform time series modeling on the communication data of each device, and extract the change trend of the communication data; Periodicity analysis: Use Fourier transform to extract the periodic characteristics of the packet size, communication frequency, data transmission direction, communication protocol, and device operation log, and identify potential periodic patterns; Outlier detection: Use a gradient-based algorithm to detect the outliers in the time series and capture the change patterns of abnormal behaviors.

5. The method according to claim 4, characterized in that, When comparing and judging the deviation degree with a preset deviation threshold, it includes: For each collected communication data Xi = [X1, X2, X3, X4, X5]; calculate the deviation degree between the communication data Xi and the corresponding behavior baseline Bi = [B1, B2, B3, B4, B5], and the deviation degree has the following relationship: where Xi represents the original value of the communication data feature of the i-th type; μ i represents the historical data mean of the i-th type of feature; σ i represents the historical data standard deviation of the i-th type of feature; d i represents the deviation degree of the i-th type of feature.

6. The method according to claim 5, characterized in that Compare and judge the deviation degree with a preset deviation threshold. When the deviation degree exceeds the deviation threshold, judge the device as abnormal and record the device information of the device Compare the total deviation degree D with the preset deviation threshold T to judge whether there is an abnormality; if D > T, it is considered that the communication data of the device is abnormal; where D is the total deviation degree.

7. The method according to claim 6, wherein It includes: When the communication data of the device is abnormal, based on the update mechanism of time series data and periodic characteristics, re-update the behavior baseline of the device.

8. The method according to claim 7, wherein When there is an abnormality and an alarm report is generated, it includes: When the total deviation degree D is less than the deviation threshold T, judge that the device is at a minor deviation level and generate a report; When the total deviation degree D is greater than or equal to the deviation threshold T, judge that the device is at a serious deviation level and generate a report.

9. A self-checking system for industrial control system network security, characterized in that, For implementing the method according to any one of claims 1-8, the system includes: A data acquisition module, configured to periodically acquire the communication data of devices in the industrial control system, including packet size, communication frequency, data transmission direction, communication protocol, and device operation log, and store historical data; A data cleaning and preprocessing module, configured to clean the acquired historical data, delete missing values and outliers, and perform necessary data correction and filling to ensure the integrity and accuracy of the data; A behavior baseline modeling module, configured to perform modeling on historical data based on the K-means clustering algorithm, extract the communication behavior characteristics of devices and form the behavior baseline of each device for subsequent anomaly detection; An anomaly detection module, configured to compare the real-time acquired communication data with the behavior baseline of the device, calculate the deviation degree of the data, and judge whether there is a device anomaly according to the preset deviation threshold; An alarm generation module, configured to generate an alarm report when a device anomaly is detected. The report includes device information, anomaly type, timestamp, and deviation degree, and send an alarm notification to the user.

10. An electronic device, characterized in that, The electronic device includes: A processor; A memory configured to store executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the method according to any one of claims 1-8.