Industrial control system anomaly detection method and device based on log, equipment and medium
By adopting log-based anomaly detection method in industrial control systems, using convolutional neural networks and graph neural networks to analyze multi-sensor data, the false alarm and missed response problems of abnormal detection in complex production processes in the prior art are solved, and more efficient and accurate abnormal detection is achieved.
Patent Information
- Application Number
- CN202510061468.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
AI Technical Summary
When facing complex production processes, the abnormality detection methods of existing industrial control systems are prone to false alarms or missed alarms, and it is difficult to effectively integrate and analyze data from multiple sensors or actuators.
The abnormal detection method of the industrial control system based on logs is adopted, and event data in the industrial control system is obtained, preprocessed and feature extraction is performed, and the relationship diagram is constructed using convolutional neural networks and graph neural networks, and the graph neural network is trained to identify abnormal events.
It improves the accuracy and efficiency of abnormal detection of industrial control systems, reduces false alarms and missed alarms, can promptly identify abnormal behaviors in the system, and ensures the stability and security of the system.
Smart Images

Figure CN120065975A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of industrial control technology, and particularly to a method, device, equipment and medium for anomaly detection of industrial control systems based on logs. Background Art
[0002] With the continuous development of industrial automation, industrial control systems (ICS) are increasingly widely used in various manufacturing and production processes. Industrial control systems ensure the efficiency, stability and safety of production by collecting, processing and controlling various physical parameters in the production process. However, with the complication of the production process, the amount of data in industrial control systems is also increasing rapidly, making it increasingly challenging to monitor and detect abnormal behaviors in the system.
[0003] To improve the intelligence level of industrial control systems, anomaly detection of industrial control systems has become an important research field. Traditional anomaly detection methods usually rely on pre-set thresholds or rules. These methods are not sensitive enough to changes in system parameters and are prone to false alarms or missed alarms when faced with complex production processes. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose a method, device, equipment and medium for anomaly detection of industrial control systems based on logs to accurately and timely detect anomalies in industrial control systems.
[0005] To achieve the above object, on the one hand, the embodiments of this application propose a method for anomaly detection of industrial control systems based on logs, and the method includes the following steps:
[0006] Obtain the first state data of each event that occurs in the industrial control system;
[0007] Preprocess the first state data of each event to obtain a plurality of standardized time series data;
[0008] Use a convolutional neural network to extract features from each of the time series data as local features;
[0009] Map each of the local features to a vector space to obtain each feature vector;
[0010] Calculate the similarity between each of the feature vectors;
[0011] Construct a relationship graph with each of the feature vectors as nodes and the similarity between each of the feature vectors as the weight of the edge corresponding to the nodes;
[0012] Use the relationship graph to train a graph neural network so that the graph neural network learns whether each event is abnormal;
[0013] Input the second state data of the event into the trained graph neural network to detect whether the event corresponding to the second state data is abnormal.
[0014] In some embodiments, obtaining the first state data of each event occurring in the industrial control system includes the following steps:
[0015] Collect event information of various sensors and various actuators in the industrial control system through the OPC AE protocol;
[0016] Collect time series data of various sensors and various actuators as the original time series data through the OPC UA protocol;
[0017] Determine the event information and the original time series data as the first state data.
[0018] In some embodiments, preprocessing the first state data of each event to obtain a plurality of standardized time series data includes the following steps:
[0019] Perform correlation processing on the event information and the corresponding original time series data in the first state data to obtain synchronous time series data;
[0020] Remove noise and outliers from each synchronous time series data, and then perform normalization processing to obtain preprocessed time series data;
[0021] Divide each preprocessed time series data into different time window data according to the occurrence time of the event and the characteristics of the industrial process as the standardized time series data.
[0022] In some embodiments, calculating the similarity between each feature vector includes the following steps:
[0023] Calculate the Euclidean distance between each feature vector as the similarity;
[0024] The calculation formula of the Euclidean distance is:
[0025]
[0026] where d(X 1 , X 2 ) represents the Euclidean distance between the feature vector X 1 and the feature vector X 2 ; k represents the serial number of the feature vector, and n represents the total number of feature vectors.
[0027] In some embodiments, training the graph neural network by using the relationship graph so that the graph neural network learns whether each of the events is abnormal includes the following steps:
[0028] Training the graph neural network by using the relationship graph and the normal data and abnormal data in the historical state data so that the graph neural network learns whether each of the events is abnormal.
[0029] In some embodiments, the method further includes the following steps:
[0030] Visualizing the relationship graph and dynamically updating the relationship graph to display the propagation path and influence range of abnormal events.
[0031] In some embodiments, visualizing the relationship graph and dynamically updating the relationship graph to display the propagation path and influence range of abnormal events includes the following steps:
[0032] Visualizing the relationship graph and dynamically updating the relationship graph by using the human-machine interface of the monitoring platform of the industrial control system to display the propagation path and influence range of abnormal events.
[0033] To achieve the above object, on the other hand, an embodiment of the present application provides a log-based industrial control system anomaly detection device, and the device includes:
[0034] A data acquisition unit, configured to acquire first state data of each event occurring in the industrial control system;
[0035] A data preprocessing unit, configured to preprocess the first state data of each event to obtain a plurality of standardized time series data;
[0036] A feature extraction unit, configured to extract features from each of the time series data by using a convolutional neural network as local features;
[0037] A feature mapping unit, configured to map each of the local features to a vector space to obtain each feature vector;
[0038] A similarity calculation unit, configured to calculate the similarity between each of the feature vectors;
[0039] A relationship graph construction unit, configured to construct a relationship graph by using each of the feature vectors as nodes and the similarity between each of the feature vectors as the weight of the edge corresponding to the nodes;
[0040] A network training unit, configured to train a graph neural network by using the relationship graph so that the graph neural network learns whether each of the events is abnormal;
[0041] An anomaly detection unit for inputting the second state data of the event into the trained graph neural network to detect whether the event corresponding to the second state data is an anomaly.
[0042] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned log-based industrial control system anomaly detection method is implemented.
[0043] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned log-based industrial control system anomaly detection method is implemented.
[0044] The embodiments of the present application at least include the following beneficial effects:
[0045] The present application can obtain the first state data of each event occurring in the industrial control system; preprocess the first state data of each event to obtain a plurality of standardized time series data; use a convolutional neural network to extract features from each time series data as local features; map each local feature to a vector space to obtain each feature vector; calculate the similarity between each feature vector; use each feature vector as a node and the similarity between each feature vector as the weight of the edge between the corresponding nodes to construct a relationship graph; use the relationship graph to train the graph neural network so that the graph neural network learns whether each event is an anomaly; input the second state data of the event into the trained graph neural network to detect whether the event corresponding to the second state data is an anomaly. By analyzing the relationship graph of each feature vector, the present application can determine the relationship between each event, and further determine the relationship between the devices corresponding to each event. Moreover, using the relationship graph to train the graph neural network can improve the recognition accuracy, and inputting the state data of the event into the trained graph neural network can obtain the detection result in time, that is, efficiently determine whether the event is abnormal. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 It is a schematic flowchart of the log-based industrial control system anomaly detection method provided by the embodiment of the present application;
[0048] Figure 2 An exemplary flowchart of the log - based industrial control system anomaly detection method provided by an embodiment of this application;
[0049] Figure 3 A structural schematic diagram of the log - based industrial control system anomaly detection device provided by an embodiment of this application;
[0050] Figure 4 A hardware structural schematic diagram of an electronic device provided by an embodiment of this application. Detailed implementation manners
[0051] To make the objectives, technical solutions and advantages of this application clearer and more understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this application. They are merely examples of devices and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0052] It can be understood that the terms "first", "second", etc. used in this application can be used in this text to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "while...", or "in response to determining".
[0053] The terms "at least one", "multiple", "each", "any one", etc. used in this application, where at least one includes one, two or more, multiple includes two or more, each refers to each one in the corresponding multiple, and any one refers to any one in the multiple.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0055] Before elaborating on the embodiments of this application in detail, some related technologies involved in the embodiments of this application are described as follows:
[0056] In recent years, with the development of artificial intelligence technology, anomaly detection methods based on machine learning and deep learning have gradually emerged, especially in the processing and analysis of time series data, showing significant advantages.
[0057] OPC (OLE for Process Control) is a data communication standard widely used in industrial automation, where OPC AE (OPC Alarms & Events) focuses on the management of alarms and events in industrial control systems. However, traditional OPC AE methods mainly focus on the simple recording and alarming of events and do not have the ability to analyze the complex relationships between events. In addition, with the popularization of industrial Internet of Things (IIoT) and intelligent manufacturing, a large amount of data generated by sensors and actuators in industrial control systems requires more intelligent processing and analysis to timely detect abnormal behaviors in the system.
[0058] Therefore, some embodiments of this application propose a method of dividing processes in an industrial control system into events through OPC AE, and combining the sensor and actuator data collected by OPC UA (OPC Unified Architecture), and processing and analyzing the event data through deep learning algorithms, thereby improving the system's ability to identify abnormal behaviors and enhancing the stability and security of the industrial control system.
[0059] Existing technical solutions:
[0060] In the prior art, the detection of abnormal behaviors in industrial control systems is usually achieved through the following several solutions:
[0061] Rule-based anomaly detection methods: Traditional anomaly detection methods mainly rely on predefined rules and thresholds. These methods judge whether there are anomalies by setting thresholds for certain key parameters, such as physical quantities like temperature and pressure. However, this method has poor robustness and accuracy when facing complex and dynamically changing production environments, and is prone to false alarms or missed alarms. In addition, with the increase in system complexity, the setting and maintenance of rules become more difficult.
[0062] Statistical analysis methods: These methods mainly include principal component analysis (PCA), time series analysis, etc., which use statistical models to detect abnormal behaviors in the system. For example, PCA identifies abnormal points by reducing the data dimension, but its ability to process non-linear data is limited and it performs poorly when facing complex industrial data.
[0063] Machine Learning-based Anomaly Detection: In recent years, machine learning techniques have been widely applied in industrial anomaly detection. Algorithms based on supervised learning and unsupervised learning, such as Support Vector Machine (SVM), Random Forest, K-means clustering, etc., predict and detect anomalies by learning historical data. However, these methods usually rely on large-scale labeled data, and when dealing with complex time-series data, their effectiveness depends on the quality of feature engineering.
[0064] Deep Learning-based Anomaly Detection: Deep learning techniques have significant advantages in processing high-dimensional and complex data, especially in the processing of time-series data. Methods such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and Long Short-Term Memory Network (LSTM) are widely used in anomaly detection of industrial control systems. For example, CNN can extract local features in time-series data, while RNN and LSTM are good at capturing long-term dependencies. However, these methods mainly focus on anomaly detection of single sensor data, and there is still a lack of modeling and analysis of the relationships between multiple sensors or actuators.
[0065] Existing anomaly detection methods have certain limitations when dealing with data in complex industrial control systems, especially in multi-sensor data fusion and multi-event correlation analysis. Therefore, some embodiments of this application propose a method that combines OPC AE and OPC UA technologies and uses Convolutional Neural Network (CNN) and Graph Neural Network (GNN) to process and analyze sensor and actuator data, so as to improve the accuracy and efficiency of anomaly detection in industrial control systems.
[0066] The following are the main disadvantages of the application of existing technical solutions in industrial control systems:
[0067] 1. Rule-based Anomaly Detection Methods:
[0068] Rely on manual setting of rules and thresholds: Such methods require experts to set rules and thresholds based on experience and the specific situation of the system, resulting in limited applicability of the rules and difficulty in adapting to the dynamically changing industrial environment.
[0069] Poor robustness to complex systems: When there are multiple variables interacting or non-linear relationships in the system, simple rules and thresholds are difficult to effectively capture abnormal behaviors, and false positives or false negatives are likely to occur.
[0070] High maintenance cost: With the complexity and changes of industrial systems, the setting and updating of rules require a large amount of time and human resources, increasing the maintenance cost of the system.
[0071] 2. Statistical Analysis Methods:
[0072] Insufficient non - linear data processing ability: Methods such as principal component analysis (PCA) are mainly applicable to linear data and perform poorly in dealing with complex non - linear relationships or high - dimensional data, making it difficult to comprehensively reflect abnormal behaviors in industrial control systems.
[0073] Sensitive to data structure: Statistical methods often assume that data has a certain distribution. However, actual industrial data may not conform to these assumptions, resulting in unsatisfactory detection effects.
[0074] Difficult feature extraction: These methods face challenges in extracting effective features from high - dimensional data, easily missing important information or introducing noise.
[0075] 3. Anomaly detection based on machine learning:
[0076] Dependent on large - scale labeled data: Supervised learning methods (such as support vector machines and random forests) usually require a large amount of labeled data for training, and obtaining accurate labeled data is often difficult and costly in industrial environments.
[0077] Complex feature engineering: When machine learning methods deal with complex time - series data, feature engineering is required to extract effective features, which requires high professional skills from engineers, and feature selection in different scenarios affects the model's performance.
[0078] Poor multi - sensor data fusion: Traditional machine learning methods are mostly used for anomaly detection of single data sources and are difficult to effectively fuse and analyze data from multiple sensors or actuators, thus unable to fully utilize the overall information of the system.
[0079] 4. Anomaly detection based on deep learning:
[0080] High computational cost: Deep learning models, especially convolutional neural networks (CNNs) and long short - term memory networks (LSTMs), have a large amount of computations during training and inference. Especially when dealing with large - scale industrial data, high - performance hardware support may be required.
[0081] Poor model interpretability: Deep learning models are often regarded as "black boxes" and it is difficult to explain the model's decision - making process. Especially in industrial applications, the interpretability of the model is very important for system maintenance and fault diagnosis.
[0082] Insufficient modeling of relationships between multiple sensors: Although deep learning can effectively process time - series data, traditional deep - learning methods usually focus on the analysis of single - sensor data and are not perfect in modeling the complex relationships between multiple sensors or actuators, unable to fully reveal the internal interactions of the system.
[0083] Generally speaking, when dealing with anomaly detection in complex industrial control systems, existing technical solutions often have deficiencies in aspects such as data fusion, model robustness, computational efficiency, and interpretability, and it is difficult to meet the requirements of modern industrial control systems for high-precision, high-efficiency, and intelligent anomaly detection. These drawbacks provide room for improvement and application prospects for the methods of some embodiments of this application.
[0084] Embodiments of this application provide a log-based anomaly detection method, device, equipment, and medium for industrial control systems. The technical solution of this application includes: obtaining first state data of each event that occurs in the industrial control system; preprocessing the first state data of each event to obtain multiple standardized time series data; using a convolutional neural network to extract features from each time series data as local features; mapping each local feature to a vector space to obtain each feature vector; calculating the similarity between each feature vector; using each feature vector as a node and the similarity between each feature vector as the weight of the edge between the corresponding nodes to construct a relationship graph; using the relationship graph to train a graph neural network so that the graph neural network learns whether each event is an anomaly; inputting the second state data of the event into the trained graph neural network to detect whether the event corresponding to the second state data is an anomaly. By analyzing the relationship graph of each feature vector, this application can determine the relationships between each event, and further determine the relationships between the devices corresponding to each event. Moreover, using the relationship graph to train the graph neural network can improve the recognition accuracy, and inputting the state data of the event into the trained graph neural network can obtain the detection result in a timely manner, that is, efficiently determine whether the event is abnormal.
[0085] Embodiments of this application provide a log-based anomaly detection method, device, equipment, and medium for industrial control systems, which relates to the field of industrial control technology. The log-based anomaly detection method, device, equipment, and medium provided by the embodiments of this application can be applied to terminals, can also be applied to servers, or can be software running on terminals or servers. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, can also be configured as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application implementing the knowledge extraction method, etc., but is not limited to the above forms.
[0086] This application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0087] Referring to Figure 1 , the embodiments of this application provide a method for detecting anomalies in industrial control systems based on logs. This method may include but is not limited to S100 to S170, specifically as follows:
[0088] S100: Obtain the first state data of each event occurring in the industrial control system.
[0089] Furthermore, S100 may include the following steps S101 to S103:
[0090] S101: Collect event information of various sensors and various actuators in the industrial control system through the OPC AE protocol;
[0091] S102: Collect time series data of various sensors and various actuators as raw time series data through the OPC UA protocol;
[0092] S103: Determine the event information and the raw time series data as the first state data.
[0093] S110: Preprocess the first state data of each event to obtain multiple standardized time series data.
[0094] Furthermore, S110 may include the following steps S111 to S113:
[0095] S111: Perform correlation processing on the event information and the corresponding raw time series data in the first state data to obtain synchronized time series data;
[0096] S112: Remove noise and outliers from each synchronized time series data, and then perform normalization processing to obtain preprocessed time series data;
[0097] S113: Divide each of the preprocessed time series data into different time window data as the standardized time series data according to the occurrence time of the event and the characteristics of the industrial process.
[0098] S120: Use a convolutional neural network to extract features from each of the time series data as local features.
[0099] S130: Map each of the local features to a vector space to obtain each feature vector.
[0100] S140: Calculate the similarity between each of the feature vectors.
[0101] Further, S140 may include the following steps S141:
[0102] S141: Calculate the Euclidean distance between each of the feature vectors as the similarity;
[0103] The calculation formula of the Euclidean distance is:
[0104]
[0105] where d(X 1 , X 2 ) represents the Euclidean distance between the feature vector X 1 and the feature vector X 2 ; k represents the serial number of the feature vector, and n represents the total number of the feature vectors.
[0106] S150: Use each of the feature vectors as a node, and use the similarity between each of the feature vectors as the weight of the edge corresponding to the nodes to construct a relationship graph.
[0107] S160: Use the relationship graph to train a graph neural network so that the graph neural network learns whether each of the events is abnormal.
[0108] Further, S160 may include the following steps S161:
[0109] S161: Use the relationship graph and the normal data and abnormal data in the historical state data to train the graph neural network so that the graph neural network learns whether each of the events is abnormal.
[0110] S170: Input the second state data of the event into the trained graph neural network to detect whether the event corresponding to the second state data is abnormal.
[0111] In some embodiments, the embodiment of the present application may further include the following steps S180:
[0112] S180: Visualize the relationship graph and dynamically update the relationship graph to display the propagation path and influence scope of abnormal events.
[0113] Furthermore, S180 may include the following step S181:
[0114] S181: Use the human - machine interface of the monitoring platform of the industrial control system to visualize the relationship graph and dynamically update the relationship graph to display the propagation path and influence scope of abnormal events.
[0115] Next, specific application examples will be combined to introduce and illustrate the solutions of the embodiments of the present application in detail.
[0116] First, the technical problems to be solved in this embodiment are described as follows:
[0117] 1. The problem of abnormal detection accuracy and efficiency in complex industrial control systems:
[0118] Existing anomaly detection methods based on rules and statistical analysis often perform poorly in complex industrial control environments, and are prone to false alarms or missed detections. Especially in multi - sensor data fusion and cross - event correlation analysis, these methods perform poorly. To address this issue, this embodiment aims to extract local features of time - series data through a convolutional neural network (CNN) and use a graph neural network (GNN) to construct and analyze the relationship graph between sensors and actuators, improving the recognition accuracy of abnormal behaviors.
[0119] 2. The problem of intelligent processing and analysis of multi - sensor and actuator data:
[0120] Traditional anomaly detection methods usually cannot effectively fuse data from multiple sensors and actuators. This embodiment proposes to collect multi - source data based on OPC UA, and through data pre - processing and deep - learning methods, perform intelligent processing and analysis on this data, thereby more comprehensively revealing the complex relationships inside the system and enhancing the perception of the system's dynamic behavior.
[0121] 3. The problem of time - series data processing and feature extraction:
[0122] The processing of time - series data is a major challenge in industrial control systems. Traditional methods rely on manual settings for feature extraction and are prone to missing key features. This embodiment automatically extracts local features of time - series data through a convolutional neural network, improving the automation and accuracy of feature extraction.
[0123] 4. The problem of interpretability and traceability of abnormal behaviors:
[0124] Deep learning models are usually regarded as "black boxes", and it is difficult to explain how the models make decisions. In this embodiment, a relationship graph of sensors and actuators is constructed through a graph neural network (GNN), providing a visual way to more intuitively display the generation path of abnormal behaviors and related data, enhancing the interpretability and traceability of the system.
[0125] This embodiment belongs to the fields of industrial automation and industrial control systems, and particularly relates to anomaly detection and diagnosis technologies in industrial control systems. In addition, the invention also relates to related technologies such as artificial intelligence, deep learning, Internet of Things (IoT), and Industrial Internet of Things (IIoT). It uses OPC UA for data acquisition and combines convolutional neural network (CNN) and graph neural network (GNN) for data analysis and processing, belonging to the category of intelligent industrial control and intelligent manufacturing.
[0126] Based on the above problems, this embodiment provides a method for dividing the processes in an industrial control system into events through OPC AE (OPC Alarms & Events). First, use OPC UA (OPC Unified Architecture) to collect data of sensors and actuators in each event and preprocess this data. Then, divide the preprocessed data into different time windows, and use a convolutional neural network (CNN) to extract local features of the time series data. By calculating the Euclidean distance, compare the similarity of data in the same event or different events to construct a relationship graph between sensors and actuators in the industrial control system. Finally, use a graph neural network (GNN) to train and process the relationship graph to discover abnormal data. This method not only improves the accuracy and efficiency of data processing but also can accurately identify abnormal behaviors in the system, ensuring the stability and security of the industrial control system.
[0127] This embodiment proposes a solution for anomaly detection in an industrial control system through OPC AE, OPC UA, convolutional neural network (CNN), and graph neural network (GNN). Exemplarily, referring to Figure 2 , the specific implementation of this embodiment includes the application of OPC AE, as follows:
[0128] Embodiment 1: Data acquisition and preprocessing based on OPC UA and OPC AE.
[0129] 1. Data acquisition:
[0130] OPC UA: Connect various sensors and actuators in the industrial control system through the OPC UA protocol to collect multi-source data generated during the operation of the system in real time. These data include physical parameters (such as temperature, pressure, flow rate) and device states (such as valve switch states, pump working states, etc.).
[0131] OPC AE: Monitor and manage alarm and event information in industrial control systems using the OPC AE protocol. When an event occurs in the system (such as over-temperature alarm, equipment failure, etc.), OPC AE will capture the detailed information of the event, including the event type, occurrence time, involved sensors or devices, etc. This event information will be synchronously processed with the sensor and actuator data collected by OPC UA.
[0132] 2. Data preprocessing:
[0133] Synchronization of events and data: Correlate the event information captured by OPC AE with the time series data collected by OPC UA to ensure that each event is synchronized with its related sensor and actuator data.
[0134] Data cleaning and integration: Clean the collected raw data to remove noise and outliers. At the same time, normalize the data to eliminate the differences in different sensor dimensions.
[0135] Time window partitioning: According to the occurrence time of events and the characteristics of industrial processes, divide the preprocessed data into different time windows. For example, the data within a certain number of seconds before and after the event occurrence can be used as a time window, thus forming a series of time series data.
[0136] W i ={x t |t∈[t i ,t i +T]};
[0137] where W i is the i-th time window, and x t is the data point at time t.
[0138] Embodiment 2: Convolutional neural network (CNN) is used for time series feature extraction.
[0139] 1. Construct a convolutional neural network:
[0140] Construct a multi-layer convolutional neural network (CNN) to extract local features from time series data. The data of each time window is processed through convolutional layers and pooling layers to gradually extract representative features.
[0141] According to the characteristics of industrial data, adjust the size and stride of the convolutional kernel to capture local patterns at different time scales.
[0142] 2. Feature mapping:
[0143] The convolutional neural network maps the extracted local features into a high-dimensional space to generate feature vectors. These feature vectors represent the patterns and changing trends of time series data within local regions.
[0144] The extracted feature vectors will serve as the basis for subsequent relationship graph construction and analysis.
[0145] Example 3: Similarity comparison based on Euclidean distance and relationship graph construction.
[0146] 1. Similarity calculation:
[0147] The similarity of data in the same event or different events is compared by calculating the Euclidean distance. Specifically, for multiple sets of data collected for the same type of event, the Euclidean distance between their feature vectors is calculated to determine whether these data have similar patterns.
[0148] If the Euclidean distance is small, it indicates that the positions of these data in the feature space are close and they may belong to similar events; if the distance is large, they may be different events or there are abnormal behaviors.
[0149]
[0150] where d(X 1 , X 2 ) represents the Euclidean distance between two feature vectors X 1 and X 2 .
[0151] 2. Relationship graph construction:
[0152] Based on the calculated similarity, the feature vectors of sensors and actuators are used as nodes, and weighted edges are used to represent the similarity between them to construct a relationship graph between sensors and actuators.
[0153] The nodes in the relationship graph represent the states of sensors or actuators, and the weights of the edges represent the similarity or correlation degree between the nodes. Through this relationship graph, the interaction relationships between various components in the industrial control system can be intuitively described.
[0154] Example 4: Event-driven graph neural network (GNN) based on OPC AE for anomaly detection and analysis.
[0155] 1. Graph neural network training:
[0156] The graph neural network (GNN) is used to train the event relationship graph constructed through OPC AE to learn the complex relationships between sensors and actuators. GNN captures the potential patterns and anomalies in the system through information transfer between nodes and edges.
[0157] During the training process, normal and abnormal samples from historical data can be introduced, and through supervised learning, the ability of the GNN to identify abnormal behaviors can be enhanced.
[0158] 2. Anomaly detection:
[0159] After the training is completed, the newly collected data is input into the GNN. By analyzing its performance in the relationship graph, it is judged whether there are abnormal behaviors. If the GNN detects that the behaviors of certain nodes or edges are significantly different from the normal patterns in the training data, it can be considered that there may be abnormal behaviors in the system.
[0160] This OPC AE event-driven based detection method can effectively focus on key events and related data, improving the accuracy and efficiency of anomaly detection.
[0161] 3. Visualization of abnormal behaviors:
[0162] To enhance the interpretability of the system, the GNN can visually display the detected abnormal behaviors. For example, through the dynamic changes of the relationship graph, the propagation path and influence range of abnormal behaviors are shown to help operators quickly locate the root cause of the problem.
[0163] Example 5: System integration and application.
[0164] 1. System integration:
[0165] Integrate the method of this example into the monitoring platform of the industrial control system to form a real-time anomaly detection and diagnosis system.
[0166] Through the interfaces with OPC UA and OPC AE servers, real-time data collection, event monitoring and processing are realized; through the integration with the HMI (Human Machine Interface) or SCADA (Supervisory Control and Data Acquisition) system, the system status and anomaly alarms are displayed in real time.
[0167] 2. Application scenarios:
[0168] The method of this example can be applied to various industrial scenarios, such as petrochemical, metallurgical, power, manufacturing and other fields, to realize intelligent detection and diagnosis of abnormal behaviors in the production process, ensuring the safety and stability of production.
[0169] By combining OPC AE for event-driven data processing and monitoring, this example not only enhances the detection ability of abnormal behaviors in the industrial control system, but also improves the efficiency and accuracy of system response. This technical solution effectively solves the problem that traditional methods are difficult to accurately detect and analyze abnormal behaviors in complex and changeable industrial environments.
[0170] In summary, the key technical features of the solution in this embodiment may include: the combined application of OPC AE and OPC UA, the feature extraction method based on CNN, and the application of GNN in the anomaly detection of industrial control systems. Each key point reflects the innovation of this embodiment in data collection, processing, and analysis technologies.
[0171] First of all, the combination of OPC AE and OPC UA is one of the basic technologies in this embodiment. In the prior art, data collection usually only relies on OPC UA, lacking comprehensive support for event-driven, which easily leads to incomplete data or delayed response. In this embodiment, by introducing OPC AE, it is possible to capture alarm and event information in the system while collecting real-time sensor data, thus realizing the synchronous processing of data and events. This combined application greatly improves the comprehensiveness and timeliness of data collection, which is difficult to achieve in the prior art.
[0172] Secondly, the feature extraction method based on CNN is an important innovation in the data processing link of this embodiment. In the prior art, feature extraction usually relies on manually designed features or simple statistical methods, which are difficult to effectively capture complex patterns in time series data. In this embodiment, the convolutional neural network (CNN) is used to automatically extract local features in the data, which can more accurately identify potential abnormal patterns in the system. This automated feature extraction method not only improves the accuracy of data processing but also reduces the need for manual intervention, having significant advantages compared with traditional methods.
[0173] Finally, the application of GNN in anomaly detection is another key point of this embodiment. Traditional anomaly detection methods mostly use simple thresholds or traditional machine learning algorithms, which are difficult to capture the complex relationships between sensors and actuators in the system. In this embodiment, by constructing a relationship graph in the industrial control system and using the graph neural network (GNN) for training and analysis, it is possible to more effectively detect abnormal behaviors, especially in complex and dynamic industrial environments. Compared with the prior art, this method not only improves the detection accuracy but also reduces the probability of false alarms and missed detections, significantly enhancing the reliability and security of the system.
[0174] In conclusion, this embodiment overcomes the deficiencies of the prior art in data collection, feature extraction, and anomaly detection by innovatively combining OPC AE, CNN, and GNN, providing a more accurate, real-time, and intelligent solution for industrial control systems.
[0175] The beneficial effects of this embodiment include:
[0176] In this embodiment, by combining the OPC UA and OPC AE protocols, comprehensive and real-time acquisition of sensor data and event information in the industrial control system is achieved. Time window partitioning and convolutional neural network (CNN) are used for feature extraction, significantly enhancing the recognition ability for complex patterns. In addition, by introducing a graph neural network (GNN) to analyze the relationship graph between sensors and actuators, the accuracy and efficiency of anomaly detection in this embodiment are improved, and the phenomena of false alarms and missed alarms are reduced. Compared with the prior art, this embodiment not only improves the accuracy and adaptability of data processing, but also enhances the integration and real-time response ability of the system, ensuring the stability and security of the industrial control system.
[0177] Referring to Figure 3 , the embodiment of the present application also provides a log-based industrial control system anomaly detection device, which can implement the above-mentioned log-based industrial control system anomaly detection method. The device includes:
[0178] A data acquisition unit, configured to acquire first status data of each event occurring in the industrial control system;
[0179] A data preprocessing unit, configured to preprocess the first status data of each of the events to obtain a plurality of standardized time series data;
[0180] A feature extraction unit, configured to extract features from each of the time series data by using a convolutional neural network as local features;
[0181] A feature mapping unit, configured to map each of the local features to a vector space to obtain each feature vector;
[0182] A similarity calculation unit, configured to calculate the similarity between each of the feature vectors;
[0183] A relationship graph construction unit, configured to construct a relationship graph with each of the feature vectors as nodes and the similarity between each of the feature vectors as the weight of the edge corresponding to the nodes;
[0184] A network training unit, configured to train a graph neural network by using the relationship graph so that the graph neural network learns whether each of the events is abnormal;
[0185] An anomaly detection unit, configured to input second status data of the event into the trained graph neural network to detect whether the event corresponding to the second status data is abnormal.
[0186] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0187] An embodiment of the present application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned log-based industrial control system anomaly detection method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0188] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0189] Please refer to Figure 4 , Figure 4 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0190] A processor 401, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0191] A memory 402, which can be implemented in forms such as a Read Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM). The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 402, and the processor 401 is used to call and execute the log-based industrial control system anomaly detection method of the embodiments of the present application;
[0192] An input / output interface 403, which is used to implement information input and output;
[0193] A communication interface 404, which is used to implement communication interaction between the device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0194] A bus 405, which transmits information between various components of the device (such as the processor 401, the memory 402, the input / output interface 403, and the communication interface 404);
[0195] Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 are communicatively connected to each other inside the device through the bus 405.
[0196] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned log-based industrial control system anomaly detection method.
[0197] It can be understood that the content in the above method embodiments is applicable to the present storage medium embodiment. The functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0198] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0199] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0200] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0201] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0202] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0203] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0204] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0205] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0206] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0207] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0208] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0209] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A log-based industrial control system anomaly detection method, characterized in that: The method comprises the following steps: Acquire first state data of each event occurring in the industrial control system; Preprocessing the first state data of each of the events to obtain a plurality of standardized time series data; Extracting features from each of the time series data as local features using a convolutional neural network; Mapping each of the local features to a vector space to obtain each feature vector; Calculating the similarity between each of the feature vectors; Using each of the feature vectors as a node and the similarity between each of the feature vectors as the weight of the edge between the corresponding nodes to construct a relationship graph; Using the relationship graph to train a graph neural network so that the graph neural network learns whether each of the events is abnormal; The second state data of the event is input into the trained graph neural network to detect whether the event corresponding to the second state data is abnormal.
2. The log-based industrial control system anomaly detection method according to claim 1, characterized in that: The method of obtaining first state data of each event occurring in the industrial control system comprises the following steps: Collect event information of various sensors and various actuators in the industrial control system through the OPC AE protocol; Collecting the time series data of the various sensors and the various actuators as raw time series data through the OPC UA protocol; The event information and the original time series data are determined as the first state data.
3. The log-based industrial control system anomaly detection method according to claim 1, characterized in that: The preprocessing of the first state data of each of the events to obtain a plurality of standardized time series data comprises the following steps: Associating the event information in the first state data with the corresponding original time series data to obtain synchronized time series data; Removing noise and abnormal points from each of the synchronized time series data, and then performing normalization processing to obtain preprocessed time series data; According to the occurrence time of the event and the characteristics of the industrial process, each of the pre-processed time series data is divided into different time window data as the standardized time series data.
4. The log-based industrial control system anomaly detection method according to claim 1, characterized in that: The calculating of the similarity between the feature vectors comprises the following steps: Calculating the Euclidean distance between each of the feature vectors as the similarity; The calculation formula of the Euclidean distance is: Wherein, d(X1, X2) represents the Euclidean distance between the feature vector X1 and the feature vector X2; k represents the sequence number of the feature vector, and n represents the total number of the feature vectors.
5. The log-based industrial control system anomaly detection method according to claim 1, characterized in that: The method of training the graph neural network using the relationship graph so that the graph neural network learns whether each of the events is abnormal includes the following steps: The graph neural network is trained using the relationship graph and normal data and abnormal data in the historical status data, so that the graph neural network learns whether each of the events is abnormal.
6. The log-based industrial control system anomaly detection method according to any one of claims 1 to 5, characterized in that: The method further comprises the following steps: The relationship graph is visualized and dynamically updated to show the propagation path and impact range of the abnormal event.
7. The log-based industrial control system anomaly detection method according to claim 6, characterized in that: The visualizing and dynamically updating the relationship graph to display the propagation path and impact scope of the abnormal event includes the following steps: The relationship diagram is visualized and dynamically updated using a human-machine interface of the monitoring platform of the industrial control system to display the propagation path and impact range of abnormal events.
8. An abnormality detection device for industrial control systems based on logs, characterized in that: The device comprises: A data acquisition unit, used to acquire first state data of each event occurring in the industrial control system; A data preprocessing unit, used for preprocessing the first state data of each of the events to obtain a plurality of standardized time series data; A feature extraction unit, used for extracting features from each of the time series data as local features using a convolutional neural network; A feature mapping unit, used for mapping each of the local features into a vector space to obtain each feature vector; A similarity calculation unit, used for calculating the similarity between each of the feature vectors; A relationship graph construction unit, used to construct a relationship graph using each of the feature vectors as a node and using the similarity between each of the feature vectors as the weight of the edge between the corresponding nodes; A network training unit, used for training a graph neural network using the relationship graph, so that the graph neural network learns whether each of the events is abnormal; An anomaly detection unit is used to input the second state data of the event into the trained graph neural network to detect whether the event corresponding to the second state data is abnormal.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the log-based industrial control system anomaly detection method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the log-based industrial control system anomaly detection method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Banking outlet safety inspection data evidence storage statistics platform based on distributed storage
CN120849511A