Systems and methods for diagnosing and monitoring anomalies in cyber-physical systems
By building an automated anomaly diagnosis and monitoring system in a cyber-physical system, and utilizing telemetry data and machine learning models, the shortcomings of existing systems in anomaly diagnosis and monitoring are addressed, enabling effective identification and handling of anomalies and improving system security and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-19
- Publication Date
- 2026-03-03
AI Technical Summary
Existing systems lack effective anomaly diagnosis and monitoring methods in cyber-physical systems (CPS), especially in telemetry data processing. They are unable to effectively classify, diagnose anomaly types, filter out unimportant anomalies, or predict anomaly development, resulting in the inability to identify and respond to potential threats in a timely manner.
By generating an automated system based on telemetry data, machine learning models are used to classify, diagnose, and monitor anomalies, including predicting anomalies and determining their characteristics. Anomaly identification module, generation module, classifier module, and diagnostic monitoring module are used, combined with a feedback interface for real-time monitoring and anomaly handling.
It enables automated diagnosis and monitoring of anomalies in CPS, and can identify, classify and predict anomalies, thereby improving the security and stability of the system and ensuring that anomalies can be identified and handled in a timely manner.
Smart Images

Figure CN116360384B_ABST
Abstract
Description
Technical Field
[0001] This invention relates generally to the field of industrial safety, and more specifically to systems and methods for diagnosing and monitoring anomalies in cyber-physical systems (CPS). Background Technology
[0002] One of the most pressing industrial safety issues is the safe operation of technical processes (TPs). The main threats to TPs include wear, tear, and failure of equipment and sub-components; unintentional or malicious errors in operational controls; and computer attacks on control systems and information systems (IS), among others.
[0003] To combat various threats, security systems are typically used to protect cyber-physical systems (CPS). Security systems can include, but are not limited to: Emergency Protection Systems (EPS), anomaly detection systems based on Automated Control Systems for a Technology Process (ACS TP), and specially constructed “external” monitoring systems for specific types of equipment and sub-components. Typically, “external” monitoring systems are not necessarily integrated with the ACS TP. It should be noted that due to certain unique aspects of CPS and TP, the aforementioned “external” systems may not always be deployable. However, even in the simplest cases where such an arrangement is possible, the deployment of such “external” monitoring systems usually occurs only at extremely critical nodes and sub-components within the enterprise due to the cost and complexity of serving them.
[0004] Compared to "external" systems, EPS can be designed during the enterprise's design phase and integrated into the ACS TP. This integration can prevent previously known critical processes from occurring. One advantage of EPS is its simplicity, its focus on a specific enterprise's production processes, and its inclusion of all design and technical solutions adopted by that enterprise. Disadvantages of EPS may include, but are not limited to: relatively slow decision-making within the system and the presence of human factors in making these decisions. Furthermore, EPS and related methods typically operate under the assumption that the Monitoring and Measuring Instrument (MMI) is functioning correctly. In practice, because MMIs experience periodic failures and have a tendency to experience temporary malfunctions, it is impossible to always ensure that MMIs operate completely without failure. Moreover, providing redundancy for all MMIs is extremely expensive and not always technically feasible.
[0005] Anomaly detection systems are typically based on telemetry technology from ACS TP (Enterprise Synchronous Processing) TP (Enterprise Synchronous Processing). Due to the completeness of this telemetry data, anomaly detection systems can simultaneously "see" the interrelationships between all TPs within an enterprise, enabling reliable anomaly detection even during MMI (Management Management System) failures. The vast amount of data provided in ACS TPs enables monitoring the entire enterprise—its physical (chemical or other) processes and the proper functioning of all monitoring systems used for these processes, which can include appropriate actions by production operators. The machine learning models used in these systems can be trained based on numerous inputs and features. Such trained models can include efficient statistical models for the proper functioning of an enterprise with a large number of variables being analyzed. These trained models can even detect minute deviations in equipment operation. In other words, anomaly detection systems can detect anomalies at an early stage.
[0006] The special architecture and interface of the anomaly detection system allow it to work in parallel with ACS TP to detect anomalies (error detection), display and localize (error isolation) the detected anomalies, and notify production operators of the detected anomalies, thereby indicating, for example, the specific process variables used to determine the anomaly.
[0007] However, existing systems using ACS TP telemetry data to identify and locate security-related anomalies and threats are not well-equipped to address the third traditional anomaly monitoring problem. More specifically, existing systems are not well-equipped to handle the technically complex issues of anomaly diagnosis itself (false diagnosis), anomaly classification based on type (category), filtering out unimportant anomalies, identifying anomalies with specific characteristics, and predicting anomaly development. Typically, the most needed types of anomaly analysis in production include assessing the hazard of certain anomalies, retrospective analysis of anomaly development, predictive assessment of one or more characteristics of anomalies, and the possibility for business operators to develop the most cost-effective corrective strategies. In most cases, specific types of analysis of previously identified anomalies are possible because ACS TP telemetry data contains exhaustive information about the operation of a particular enterprise, the progress of all its physical, chemical, and other processes, and comprehensive information about control processes. However, this telemetry information typically contains only raw, unprocessed, and unlabeled data.
[0008] Therefore, there is a need for automated systems that can effectively diagnose and monitor previously identified anomalies in CPS based on telemetry data. This need is urgent for all CPS systems that include any of the following: MMI, actuators, or monitoring systems. Summary of the Invention
[0009] Systems and methods for creating automated systems based on telemetry data for diagnosing and monitoring anomalies previously identified in CPS are disclosed.
[0010] Advantageously, the disclosed method automatically diagnoses and monitors anomalies in the CPS by classifying previously discovered anomalies, diagnosing anomalies in each category, and subsequently monitoring the CPS to identify anomalies in each category.
[0011] In one aspect, a method for diagnosing and monitoring anomalies in a cyber-physical system (CPS) includes obtaining information related to anomalies identified in the CPS. The obtained information includes at least one value of one or more CPS variables. One or more classification features of the anomalies identified in the CPS are generated based on the obtained information. Based on the generated classification features, the anomalies identified in the CPS are classified into two or more anomaly categories. Each of the two or more anomaly categories is associated with one or more anomalous characteristics. Anomaly diagnosis is performed in each of the two or more anomaly categories by calculating the value of the anomalous characteristic associated with each of the two or more anomaly categories. Anomalies in each of the two or more anomaly categories are monitored based on the calculated values of the anomalous characteristics associated with each of the two or more anomaly categories.
[0012] In one aspect, monitoring CPS to identify anomalies further includes: predicting at least one value of one or more CPS variables; determining a total prediction error based on the predicted values of the one or more CPS variables; and identifying an anomaly if the determined prediction error exceeds a predetermined threshold.
[0013] In one aspect, monitoring CPS to identify anomalies also includes identifying anomalies by applying a trained machine learning model to at least one value of the one or more CPS variables.
[0014] In one aspect, monitoring CPS to identify anomalies further includes: determining whether at least one value of the one or more CPS variables is outside the limits of a previously specified range of values for the respective CPS variable; and identifying an anomaly in response to determining that the value of at least one of the one or more CPS variables is outside the limits of a previously specified range of values for the respective CPS variable.
[0015] In one aspect, the information obtained also includes at least one of the following: the time interval between observations of the detected anomaly, the contribution of each of the one or more CPS variables to the detected anomaly, information about the detection method of the detected anomaly, and the value of the one or more CPS variables at each moment of the observation time interval.
[0016] In one aspect, for each of the one or more CPS variables, the information obtained also includes: the time series of the corresponding CPS variable values; the current magnitude of the deviation of the predicted CPS variable values from the actual CPS variable values; and the smoothed value of the deviation of the predicted CPS variable values from the actual CPS variable values.
[0017] In one aspect, the values of the one or more CPS variables include at least one of the following: a measurement value from a data transmitter; a value of a manipulated variable of an actuator; a setpoint of the actuator; one or more values of the input signal of a proportional-integral-derivative (PID) controller; and a value of the output signal of the PID controller.
[0018] In one aspect, the one or more categorical features are generated by assigning one or more CPS variables to each of the one or more categorical features.
[0019] In one aspect, the method further includes transforming the values before assigning values to one or more CPS variables for the corresponding categorical features.
[0020] In one aspect, the one or more classification features are generated based on feedback from the CPS operator. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate one or more exemplary aspects of the invention, and together with the detailed embodiments, are used to explain the principles and implementations of these exemplary aspects.
[0022] Figure 1a shows a schematic diagram of an exemplary technical system.
[0023] Figure 1b This illustrates a specific example of how a technical system is implemented.
[0024] Figure 1c This is a diagram illustrating a possible variation of the organizational structure of the Internet of Things (IoT) on an example of a portable device.
[0025] Figure 1d A block diagram showing one possible configuration of the device's data transmitter is presented.
[0026] Figure 2 This is a schematic diagram illustrating a CPS with defined characteristics and an example of a system for detecting, classifying, and monitoring anomalies.
[0027] Figure 3 This is a schematic diagram of a system used for diagnosing and monitoring anomalies in CPS.
[0028] Figure 4 This is an example of an exception determination module.
[0029] Figure 5 This is an example of a module used for diagnosing and monitoring anomalies.
[0030] Figure 6 This is a flowchart illustrating an example method for diagnosing and monitoring anomalies in CPS.
[0031] Figure 7 Examples of computer systems on which aspects of the systems and methods disclosed herein can be implemented are shown. Detailed Implementation
[0032] This document describes exemplary aspects in the context of systems, methods, and computer program products for diagnosing and monitoring anomalies in cyber-physical systems. Those skilled in the art will recognize that the following description is merely illustrative and not intended to be limiting in any way. Other aspects will readily occur to those skilled in the art upon appreciating the advantages of the invention. Implementations of exemplary aspects as illustrated in the accompanying drawings will now be referenced in detail. Throughout the drawings and the following description, the same reference numerals will be used as much as possible to refer to the same or similar items.
[0033] Glossary: This document defines a number of terms that will be used to describe various aspects of the invention.
[0034] A controlled object is a technical object to which external actions (control and / or disturbance actions) are applied to change its state. In particular, a controlled object can be a device (e.g., an electric motor) or a technical process (or a part thereof).
[0035] Technical process (TP) — The process of producing materials, which includes the continuous change of the state of a material entity (e.g., a work object).
[0036] A control loop is the material entity and control function required to automatically adjust the values of metering process variables to achieve the desired setpoint values. A control loop may include, but is not limited to, data transmitters and sensors, controllers, and actuators.
[0037] Process Variable (PV) — The current measurement value of a specific part of TP that is being observed or monitored. For example, the measurement value of a data transmitter can be a process variable.
[0038] Setpoint—the value of a process variable that is to be maintained.
[0039] Manipulated variable (MV) — A variable that is adjusted to maintain the value of a process variable at a setpoint level.
[0040] External action—a method of changing the state of a component (e.g., a component of a technical system (TS)) by applying the action in a certain direction. An external action can be transmitted from one component of the TS to another component of the TS in the form of a signal.
[0041] The state of a controlled object—all its fundamental attributes—is represented by variables indicating its state to be changed or maintained under the influence of external actions, including but not limited to control actions applied to parts of the control subsystem. These state variables are one or more numerical values characterizing the object's fundamental attributes. These state variables can be numerical values of physical quantities.
[0042] Formal status of a controlled object—the status of a controlled object corresponding to process diagrams and other process documents (if the controlled object involves a process flow diagram, TP) or movement travel (if the controlled object involves a device).
[0043] Control action—A goal-oriented (the goal of the action is the action applied to the control body of the control subsystem on the controlled object), legal (specified by TP) external action, thereby causing a change in the state of the controlled object or maintaining the state of the controlled object.
[0044] A control subject is a device that applies a control action to a controlled object or transmits the control action to another control subject for conversion before applying it directly to the controlled object.
[0045] The state of a controlling subject—all its fundamental attributes—is represented by variables representing its state, which are to be changed or maintained under the influence of external actions. These state variables are one or more numerical values characterizing the subject's fundamental attributes. These state variables can be numerical values of physical quantities.
[0046] The basic attributes of a control subject (or, also known as basic state variables) are attributes that directly affect the state of the controlled object. The basic attributes of a controlled object are attributes that directly affect monitored factors (such as, but not limited to, accuracy, safety, and effectiveness) of the operation of the control system (TS). For example, basic attributes may include the compliance of cutting conditions with formally specified conditions, the movement of a train according to its journey, and the maintenance of reactor temperature within permissible limits. Based on the monitored factors, variables representing the state of the controlled object and the associated state variables of the control subject that applies control actions to the controlled object are selected.
[0047] Multi-level control subsystem—involving all control entities at multiple levels.
[0048] Cyber-physical systems (CPS) are a concept in information technology that represents the integration of computing resources in physical processes. In a CPS system, data transmitters, devices, and computer systems are connected along the entire value creation chain, extending beyond the framework of a single enterprise or company. These systems interact with each other via standard Internet protocols to predict, automatically adjust, and adapt to change. Examples of cyber-physical systems include, but are not limited to, technological systems, the Internet of Things (IoT) (including portable devices), and the Industrial Internet of Things (IIoT).
[0049] The Internet of Things (IoT) is a network of computers consisting of physical objects (“things”) equipped with embedded networking technology for interacting with each other or with the outside world. IoT can include, but is not limited to, portable devices, electronic systems in vehicles, smart cars, smart cities, and industrial systems.
[0050] The Industrial Internet of Things (IIoT) – devices and platforms that extend analytics, connect to the internet, and process data acquired from connected devices. IIoT devices can be as diverse as possible – ranging from small weather data transmitters to complex industrial robots. While the term "industrial" conjures images of associated structures such as warehouses, shipyards, and factories, IIoT technology has enormous potential for use in the widest range of sectors, including but not limited to agriculture, healthcare, financial services, retail, and advertising. The Industrial Internet of Things is a subcategory of the Internet of Things (IoT).
[0051] A technical system (TS) is a multi-level control subsystem comprising all functionally interconnected control entities and controlled objects (TPs or devices) that enable changes in the state of the controlled objects by altering the state of the control entities. The architecture of a technical system consists of its basic components (the interconnected control entities and controlled objects of the multi-level control subsystems) and the links between these components. When the controlled object in a technical system is a technical process, the ultimate goal of control is to change the state of the work object (raw material, machined blank, etc.) by changing the state of the controlled object. When the controlled object in a technical system is equipment, the ultimate goal of control is to change the state of that equipment (e.g., a vehicle, a spacecraft, etc.). The functional interdependence of components in a TS refers to the interdependence between the states of these components. These components may not even have an immediate physical link. For example, there may be no physical link between an actuator and a technical operation. For example, cutting speed is functionally related to the spindle's rotational speed, even if these state variables are physically unrelated.
[0052] Computer attacks (also known as cyberattacks) are deliberate actions, using hardware and software, against computer systems and computer / telecommunications networks to breach information security in those systems and networks.
[0053] Figure 1aA schematic diagram of an exemplary technical system (TS) 100 is shown. In one aspect, the components of the TS may include, but are not limited to: a controlled object 110a; a control body 110b; a multi-level control subsystem 120; horizontal links 130a and vertical links 130b. The control body 110b is grouped through levels 140.
[0054] Figure 1b A specific example of the implementation of the technical system 100' is schematically shown. The controlled object 110a' may include, but is not limited to, a TP or device. Control actions can be applied to the controlled object 110a', which can be executed and implemented by an Automated Control System (ACS) 120'. In one aspect, the ACS 120' may include three levels 140', each level 140' including a control body 110b', which is horizontally linked (links within the level). Figure 1b (Not shown) are interconnected and vertically linked via vertical links 130b' (links between levels). These interconnections can be functional. In other words, in general, a state change in a control subject 110b' at one level can cause state changes in other control subjects 110b' connected to that control subject 110b' at that level and at other levels. Information about state changes in control subjects 110b' can be transmitted as signals along the horizontal and vertical links established between control subjects 110b'. For example, information about a specific control subject 110b''' can be an external action related to other control subjects 110b'. Levels 140' in ACS 120' can be defined according to the purpose of the control subject 110b'. The number of levels can vary depending on the complexity of ACS 120'. Simple systems may contain one or more lower levels. Wired networks, wireless networks, and integrated microcircuits can be used for the physical links between the components (110a', 110b') of the TS and the subsystems of the TS 100'. Ethernet, Industrial Ethernet, and Industrial Networks can be used for logical links between TS components (110a', 110b') and TS 100' subsystems. Different types and standards that can be used by Industrial Networks and Protocols include, but are not limited to: Profibus, FIP, Controlnet, Interbus-S, DeviceNet, P-NET, WorldFIP, LongWork, Modbus, etc.
[0055] Higher levels (Supervisory Control and Data Acquisition (SCADA) level) can be levels controlled by dispatchers and operators. Higher levels can include, but are not limited to, at least the following control entities 110b': controllers, control computers, and human-machine interfaces (HMIs). It should be noted that... Figure 1b The SCADA within a single control unit is shown. Higher levels can be configured to track the status of the TS's components (110a', 110b'), acquire and store information about the status of the TS's components (110a', 110b'), and correct that status if necessary.
[0056] The intermediate level (CONTROL level) can be a controller level. The intermediate level can include, but is not limited to, at least the following control entities 110b': programmable logic controllers (PLCs), counters, relays, and regulators. A PLC-type control entity 110b' can be configured to obtain information about the status of the controlled object 110a' from monitoring and measuring instrument-type control entities 110b' and data transmitter-type control entities 110b'. A PLC-type control entity 110b' can also be configured to create control actions according to a programmable control algorithm used by an actuator-type control entity 110b'. An actuator can be configured to directly implement a given control action (apply it to the controlled object) at a lower level. An actuator can be a component of an execution device (facility). For example, but not limited to, a proportional-integral-derivative controller or a PID controller, the regulator can be a device with feedback in the control loop.
[0057] The lower level (input / output level) can be a level including, but not limited to, control subjects 110b' such as, but not limited to, data transmitters and sensors, monitoring and measuring instruments (MMIs), actuators, etc., which monitor the state of the controlled object 110a'. Actuators can be configured to act directly on the state of the controlled object 110a' to align it with the formal state. The formal state can include, for example, states corresponding to a technical work sequence, process diagram, or other process document (in the case of a TP) or movement stroke (in the case of a device). At this lower level, signals from a data transmitter-type control subject 110b' can be coupled with inputs from intermediate-level control subjects 110b'. Control actions prepared by a PLC-type control subject 110b' can be coupled with actuator-type control subjects 110b' that implement these control actions. An actuator can be a component of an execution device. The execution device can be configured to move an adjusting element based on signals from a regulator or control device. The execution device is the last link in the automatic control chain. Typically, the execution device can include, but is not limited to, the following units:
[0058] • Amplification equipment (contactors, frequency converters, amplifiers, etc.);
[0059] • Actuators (electric, pneumatic, or hydraulic actuators) with feedback elements (detectors for output shaft position, signal transmitters for end position, manual drives, etc.);
[0060] • Regulating elements (gates, valves, sliders, etc.).
[0061] The design of the actuating device can vary depending on the application conditions. The actuator and regulating elements are usually located in the basic unit of the actuating device.
[0062] In a particular example, the execution device may include an actuator.
[0063] It should be noted that the planning and control tasks of an enterprise can be handled by ACSE (Automatic Control System for an Enterprise) 120a', and ACSE 120a' can be a part of ACS 120'.
[0064] Figure 1c This is a diagram illustrating one possible variation of the organizational structure of an example of the Internet of Things (IoT) based on portable devices. Figure 1c The system shown may include, but is not limited to, a different set of computer devices 151 for the user. User devices 151 may include, but are not limited to, smartphones 152, tablets 153, laptops 154, portable devices such as augmented reality glasses 155, “smart” watches 156, etc. User devices 151 may include a different set of data transmitters 157a-157n, such as, but not limited to, a heart rate monitor 2001 and a pedometer 2003.
[0065] It should be noted that data transmitters 157a-157n can exist on a single user device 151 or on multiple devices. Furthermore, some data transmitters 157a-157n can exist simultaneously on multiple user devices 151. Some data transmitters 157a-157n can exist as multiple units. For example, a Bluetooth module can exist on all user devices 151, while a smartphone 152 can include two or more microphones required for noise suppression and determining the distance to a sound source.
[0066] Figure 1d A block diagram illustrating one possible configuration of the data transmitter of device 151 is presented. For example, the following items may exist in data transmitters 157a-157n:
[0067] Heart rhythm monitor (heartbeat transmitter) 2001, which can be configured to determine a user's pulse. In one aspect, heart rhythm monitor 2001 may include electrodes and can measure an electrocardiogram;
[0068] Blood oxygen saturation detector 2002;
[0069] Pedometer 2003;
[0070] Fingerprint detector 2004;
[0071] Gesture Detector 2005 can be configured to recognize user gestures;
[0072] Camera 2006, such as a camera pointing around the user and a camera pointing towards the user's eyes, the camera pointing towards the user's eyes can be configured to determine the movement of the user's eyes and verify the user's identity based on the iris or retina of the eyes;
[0073] User body temperature detector 2007 (e.g., a body temperature detector that comes into direct contact with the user's body, or a non-contact body temperature detector);
[0074] Microphone 2008;
[0075] Ultraviolet Radiation Detector 2009;
[0076] Positioning system receiver 2010, such as but not limited to: GPS, GLONASS, BeiDou, Galileo, DORIS, IRNSS, QZSS, or other receivers;
[0077] GSM module 2011;
[0078] Bluetooth module 2012;
[0079] Wi-Fi module 2013;
[0080] Room temperature detector 2014;
[0081] The barometer 2015 can be configured to measure atmospheric pressure and determine the altitude above sea level based on atmospheric pressure.
[0082] A geomagnetic sensor 2016 (e.g., an electronic compass) can be configured to determine a basic orientation and azimuth angle;
[0083] Humidity detector 2017;
[0084] The luminance detector 2018 can be configured to determine color temperature and luminance levels;
[0085] The proximity detector 2019 can be configured to determine the distance to various objects located nearby;
[0086] Image Depth Detector 2020, which can be configured to obtain three-dimensional images of space;
[0087] Accelerometer 2021, which can be configured to measure acceleration in space;
[0088] The gyroscope 2022 can be configured to determine position in space;
[0089] The Hall detector 2023 (magnetic field detector) can be configured to determine the strength of a magnetic field;
[0090] The radiometer / radiometer 2024 can be configured to determine radiation levels;
[0091] NFC module 2025;
[0092] LTE module 2026.
[0093] Figure 2 This is a schematic diagram illustrating an example of a cyber-physical system 200 with certain characteristics and a system 201 for detecting, classifying, and monitoring anomalies. Figure 2 The CPS 200 is shown in a simplified form. Examples of the CPS 200 may include the aforementioned Technology System (TS) 100 (see [link to documentation]). Figures 1a-1b ), Internet of Things (see) Figures 1c-1d (and the Industrial Internet of Things). For illustrative purposes only, TS will be discussed here as a basic example of CPS 200. As mentioned above... Figures 1a-1b The CPS200 may include, but is not limited to, a set of control entities such as data transmitters, actuators, and PID controllers. For example, unprocessed data from these control entities may be sent to the PLC via analog signals. The PLC may be configured to process the data and convert it into digital form—values of CPS variables. CPS variables may include, but are not limited to, CPS process variables (i.e., telemetry data from the CPS 200). The values of the CPS variables may be sent to SCADA system 110b' and system 201 discussed herein.
[0094] The variables in a CPS can be numerical characteristics of the control entity (data transmitter, actuator, and PID controller). Therefore, the values of CPS variables can include, but are not limited to, at least one of the following: the measured value (reading) of the data transmitter; the value of the manipulated variable of the actuator; the setpoint of the actuator; the value of the input signal of the proportional-integral-derivative (PID) controller; the value of the output signal of the PID controller; and the values of other process variables of the CPS.
[0095] In one respect, the value of a CPS variable can take the form of the following set of values: [identifier (variable name), time, value]. For example, if the CPS variable is a temperature detector, then the value of the CPS variable can be represented as follows: [temperature detector, 01.01.2022 10:00:00, 99℃].
[0096] The values of the variables in the CPS 200 can be used by the anomaly determination module 202, which can be configured to determine anomalies in the CPS 200. Anomalies in the CPS 200 can be events that deviate from the normal values of one or more variables in the CPS. For example, anomalies can occur in the CPS 200 due to computer attacks, incorrect or unauthorized human intervention in the TS or TP operation, errors or deviations in the technical process (including errors or deviations involving changes in operating conditions), transitions from control loops to manual mode, incorrect readings from data transmitters, and other well-known causes. Information about anomalies found in the CPS 200 can be sent to the system 300 for diagnosing and monitoring anomalies.
[0097] Figure 3 This is a schematic diagram of system 300 used for diagnosing and monitoring anomalies in CPS 200. System 300 can be a computer system, for example... Figure 7 The system shown. System 300 may include, but is not limited to, hardware processor 21 and memory 22 (such as...). Figure 7 (As shown). System 300 may include functional and / or hardware modules and devices, which may in turn contain instructions for execution on hardware processor 21. The following describes various aspects of the aforementioned modules of system 300.
[0098] System 300 may include an aggregation module 302 configured to collect information about anomalies in CPS 200 that have been identified by anomaly determination module 301. Examples of anomaly determination module 301, and particularly modules 401-405, are described below. Figure 4 Presented in the middle.
[0099] Description of the anomaly determination module 301.
[0100] The anomaly detection module 301 can be configured to determine anomalies 401 in the CPS by predicting the values of variables in the CPS (“CPS variables”) and subsequently determining the total prediction error of the CPS variables. The anomaly detection module 301 can also be configured to detect anomalies in the CPS 200 if the total prediction error is greater than a threshold. Furthermore, the anomaly detection module 301 can determine the contribution of the CPS variables to the total prediction error as the contribution of the prediction error of the corresponding variable in the CPS to the total prediction error.
[0101] Figure 4 This is an example of an exception determination module.
[0102] The anomaly detection module 301 may include a module for identifying anomalies using a trained basic machine learning module (hereinafter referred to as the basic model module) 402 based on the variable values of the CPS. The basic model module 402 for anomaly identification can be trained using data from teaching samples, regardless of whether it includes known anomalies in the CPS 200 and the values of the CPS variables within a given time period. To improve the quality of the basic model 402, test samples and validation samples can be used to test and validate the trained basic model 402, respectively. Test samples and validation samples may include, but are not limited to, known anomalies in the CPS 200 and the values of the CPS variables within a given time period prior to the known anomalies, but may differ from the teaching samples.
[0103] On the other hand, the anomaly determination module 301 may include a rule-based determination module 403, which can be configured to determine anomalies using rules. These rules may be pre-defined and obtained from the operator 330 of the CPS via the feedback interface 320. These rules may contain conditions applicable to variable values of the CPS, and when these conditions are met, an anomaly is determined to exist.
[0104] In another aspect, the anomaly determination module 301 may include a limit-based determination module 404, which may be configured to determine that an anomaly exists when the value of at least one variable of the CPS exceeds a value range previously established for that variable for the CPS. These value ranges may be calculated based on the characteristics of the CPS 200 or the values in the file, or may be obtained from the operator 330 of the CPS via the feedback interface 320.
[0105] On the other hand, the anomaly determination module 301 may include a determination module 405 based on a set of methods, which may be configured to determine the presence of an anomaly in the CPS 200 by averaging the results of the work of the set of methods 405 (e.g., logic may be combined and applied to the results of the work of different methods) using a set of methods that may be implemented by modules 401-404.
[0106] On another front, the anomaly detection module 301 may include a graphical interface system for manual anomaly detection by the CPS operator 330, with related information transmitted via the feedback interface 320.
[0107] In one aspect, information about anomalies in CPS 200 may include, but is not limited to, the following descriptions of the anomalies: the time interval for observing the anomalies, the contribution of each variable in the CPS to the anomaly, information about the method for identifying the anomalies, and the values of the CPS variables at each time interval. In another aspect, information about anomalies in CPS 200 may also include, for each CPS variable, at least one of the following: a time series of values, the current magnitude of the deviation of the predicted value from the actual value, and a smoothed value of the deviation of the predicted value from the actual value. In yet another aspect, information about anomalies may include information about the module (method) used to identify the anomalies.
[0108] Description of abnormal database 310.
[0109] Return to reference Figure 3 The identified anomaly information may include a list of CPS variables, the values of CPS variables within a given time interval, and the aforementioned additional information. Aggregation module 302 can store this information for each anomaly in an anomaly database 310. Anomaly database 310 may be contained in memory 22. Memory 22 includes permanent storage device (ROM) 24 and random access memory (RAM) 25 (e.g., ...). Figure 7 (As shown). Therefore, the exception database 310 can be contained in both ROM 24 and RAM 25. The operator 330 of the CPS can also access the exception database 310 through the feedback interface 320 to obtain complete and current information about exceptions in the CPS 200.
[0110] Different types of databases can be used as the exception database 310, including but not limited to: hierarchical databases (IMS, TDMS, System 2000), network-based databases (Cerebrum, Cronospro, DBVist), relational databases (DB2, Informix, Microsoft SQL Server), object-oriented databases (Jasmine, Versant, POET), object-relational databases (Oracle, PostgreSQL, FirstSQL / J), functional databases, time-series databases (InfluxDB), and so on. Furthermore, the exception database 310 can be implemented as a list or data archive of exceptions stored in a file in storage 22.
[0111] Describe the classification features.
[0112] System 300 may also include a generation module 303, which is connected to the aggregation module 302 and configured to form classification features of the identified anomalies based on the collected information. These classification features may be stored in an anomaly database 310.
[0113] In one aspect, the generation module 303 can be configured to form categorical features by assigning the values of the CPS variables in their initial form, or in a transformed form, or as a result of a function of the variables applied to the CPS. For example, the sample mean or sample variance of the CPS variables can be selected as categorical features. In another aspect, categorical features can be obtained by the generation module 303 as the result of a Fourier analysis of the CPS variables. In yet another aspect, categorical features can be obtained as the result of a principal component analysis (PCA) of the variables applied to the CPS. In yet another aspect, information about one of the anomaly determination modules 301 used to identify anomalies can also be selected as categorical features. Therefore, a list of categorical features can be presented by the generation module 302 in the form of a vector of categorical feature values. This set of categorical features and the method of forming them can be specified in advance or obtained from the CPS operator 330 through the feedback interface 320. In particular, for a specific type of CPS 200 and the processes occurring in that type of CPS, the set of categorical features and their numerical computation techniques can be known in advance. For example, in the case of diagnosing and monitoring anomalies in the walls of oil pipelines using magnetic flaw detection, the size of the echo from the defect, the maximum value of the echo, and the shape of the echo signal in the diagnostic data can be classification features.
[0114] Describe the classification steps.
[0115] System 300 may also include a classifier module 304 connected to generation module 303. Classifier module 304 may further include a teaching module 305 and a classification module 306. Teaching module 305 may be configured to adjust classification rules based on classification features from an anomaly database 310. In one aspect, these classification rules may include a supervised machine learning model or an unsupervised machine learning model (e.g., a clustering model). In these aspects, adjusting the classification rules may involve forming training samples, including values of classification features for historical time periods containing the time intervals in which anomalies were observed. Furthermore, test samples and validation samples may also be formed by teaching module 305, similarly containing values of classification features for historical time periods. These samples may be stored in anomaly database 310.
[0116] In one respect, when classifying anomalies using an unsupervised model (i.e., by using a clustering model), the classifier module 304 may select one or more of the following methods as the clustering model:
[0117] Hierarchical clustering;
[0118] Density-based noise spatial clustering (DBSCAN);
[0119] Neural Gas Growth Algorithm (GNG);
[0120] An algorithm for point sorting to identify cluster structures (OPTICS).
[0121] It should be noted that the above method can compare anomalies and determine the multiple anomaly formation categories already identified by module 301 based on different anomalies. It should also be noted that other clustering methods known in the art can be used, such as, but not limited to, the K-means (K-Means) method.
[0122] On the other hand, anomalies can be classified into predetermined categories using labeled historical samples. In other words, anomalies can be classified using a classification model (supervised learning model). As the classification model, the classification module 306 can select any machine learning model known in the art for classification, including but not limited to logistic regression, neural networks, decision trees, gradient boosting on decision trees, reference vector methods, etc. The list of predetermined categories can be obtained from the operator 330 of the CPS through the feedback interface 320. Alternatively, the list of predetermined categories can be obtained from a clustering method by the classification module 306.
[0123] In addition, the classification module 306 can use a set of two or more clustering or classification models to make a decision by voting on the individual models in the set.
[0124] The classification module 306 can also be configured to adjust the classification rules to classify identified anomalies into at least two categories based on classification features. The classification module 306 can perform the classification of anomalies from the anomaly database 310 and the anomaly determination module 301 in real-time or streaming mode. In other words, the generation module 303 can also be configured to operate in streaming mode, thereby processing all incoming anomalies sequentially or in parallel.
[0125] The generated exception categories can be stored in the exception database 310, and when needed, the exception categories can be sent to the CPS operator 330 through the feedback interface 320.
[0126] System 300 may also include a diagnostic and monitoring module 307 connected to an anomaly database 310, a classifier module 304, and a feedback interface 320. The diagnostic and monitoring module 307 can be configured to obtain information about anomalies and anomaly categories. Subsequently, the diagnostic and monitoring module 307 can perform diagnostics independently and then monitor anomalies for each anomaly category. The following will combine... Figure 5 The diagnostics and monitoring module 307 will be discussed in more detail.
[0127] Describe the diagnostics and monitoring module 307.
[0128] Figure 5This is an example of an anomaly diagnosis and monitoring module. Therefore, the diagnosis and monitoring module 307 may include a diagnosis module 501, a filtering module 502, a retrospective analysis module 503, a predictive analysis module 504, a flow analysis module 505, and a module 506 for making recommendations on anomaly handling.
[0129] Monitoring of each category of anomalies may involve performing at least one of the following types of analysis at a given frequency, based on data of characteristic values for each anomaly category: retrospective analysis by module 503, predictive analysis by module 504, and stream analysis by module 505. It should be noted that the characteristics of each anomaly category are equivalent to the anomaly characteristics associated with each anomaly category. The aforementioned monitoring frequency may be predetermined or indicated by the CPS operator 330 via feedback interface 320. In one aspect, the monitoring frequency can be determined by the time interval at which monitoring is performed. For example, the monitoring frequency could be hourly or daily. In another aspect, the monitoring frequency can be determined by the conditions under which monitoring occurs. For example, the diagnostic and monitoring module 307 may determine the monitoring frequency when a predetermined number of new anomalies occur. Therefore, the monitoring results (information about the anomalies) for each category of anomalies can be complete and up-to-date.
[0130] The diagnostic module 501 can be configured to calculate a specific set of characteristics for each category of anomalies. In certain aspects, this set of characteristics can be calculated based on the characteristics of the technical processes of the CPS 200, the composition of the CPS 200's devices and sub-components, and industry standards for a given CPS 200. For example, the values of these characteristics can be calculated by the diagnostic module 501 by assigning at least the following values: CPS variable values determined for a given anomaly category, derivatives of these variables of the CPS 200, statistical characteristics of the CPS variable values, numerical values of frequency analysis of the CPS variables, etc. Therefore, preventing potential computer attacks on the technical processes of the CPS 200 is an urgent technical problem for a wide range of CPS 200 categories. A non-limiting example of such a computer attack is to spoof or replace data of certain CPS variables to disrupt feedback loops in the control circuitry of the CPS 200 and subsequently cause potential damage to the CPS 200's devices and sub-components. To prevent such computer attacks, predictive methods for anomaly determination can be used. Selected characteristics may include the "stickiness" (values repeating over time) of certain CPS variables at the same location, and the duration of this "stickiness." Alternatively, data substitution can be accomplished by repeatedly reviewing the same portion of the data. In this case, selectable characteristics may include, but are not limited to, window statistics of the signal, i.e., sampling points of the signal (e.g., sample mean), variance, autocorrelation, and cross-correlation points.
[0131] Therefore, in certain aspects, the diagnostic module 501 can calculate the value of a characteristic of an anomaly category based on the CPS variable. For example, the diagnostic module 501 can calculate the value of at least one of the following characteristics: the minimum and maximum values of the CPS variable, the statistical characteristics of the CPS variable (especially the sample mean and sample variance), the presence and characteristics of trends in the behavior of the CPS variable, the spectral characteristics of the CPS variable (e.g., the coefficients of the Fourier transform, the presence of certain vibration modes), and other characteristics.
[0132] Furthermore, characteristic values may include, but are not limited to, calculated values of the critical level of a specific anomaly, the frequency of occurrence of an anomaly of a given category, the periodicity of such occurrence, the appearance of certain vibration modes at previously known or unknown frequencies, etc. As used herein, the critical level of an anomaly is defined by the value of a given characteristic of the anomaly or the value of a given set of characteristics exceeding one or more predetermined levels, where a negative process may occur in the TP of the CPS. Therefore, for rotating equipment, including critical sub-components such as circulating pumps, anomalies involving the vibration of the rotating mechanism are characteristics, and these characteristics can typically be diagnosed based on data from vibration velocity and vibration acceleration sensors. When dealing with such anomaly characteristics, the diagnostic module 501 can select the maximum window value of the vibration analysis data. Furthermore, this set of characteristics can be extended to include characteristics such as, but not limited to, windowed Fourier transforms and specified mode ranges for monitoring. Such an extended set of characteristics enables the diagnostic module 501 to thoroughly diagnose and monitor vibration anomalies of the circulating pump and detect the emergence of new parasitic vibration modes at an early stage, thereby predicting the development of such anomalies over time.
[0133] On the other hand, for example, the set of characteristics of each category of anomalies can be obtained from the operator 330 of the CPS through the feedback interface 320.
[0134] The filtering module 502 can be configured to create rules for filtering anomalies of a certain category based on the diagnostic results of the diagnostic module 501. For example, the filtering module 502 can create filtering rules where all anomalies with characteristics obtained by module 501 and exceeding a specific range of predetermined values can pass. For instance, all anomalies from a category with a low critical level can be filtered out, i.e., passed. Another example of anomalies that may need to be filtered out is incorrectly identified anomalies or anomalies specifically noted by the CPS operator 330, as well as anomalies involving legitimate human intervention in enterprise processes. Such anomalies can be picked out by the CPS operator 330 or by the filtering module 502 in an automatic manner (e.g., based on data from the PID controller setpoint at its point of sudden change).
[0135] The retrospective analysis module 503 can be configured to perform retrospective analysis on the characteristics of anomalies involving individual devices and anomaly categories (as obtained by the diagnostic module 501). The retrospective analysis module 503 can select values of characteristics for anomaly categories in CPS200 from the anomaly database 310 for a single device or subcomponent selected by the CPS operator 330 within a given observation period. The values of characteristics for each anomaly category calculated by the diagnostic module 501 can then be used by the retrospective analysis module 503 to perform analysis. For each anomaly category, a vector (the set of values) of the anomaly's characteristics over a specific historical time interval can be analyzed by applying a machine learning model for retrospective analysis. The machine learning model used by the retrospective analysis module 503 can include, but is not limited to, regression analysis models and interpolation models. These models can receive the values of the anomaly's characteristics over a given historical time interval as their input. The results of the retrospective analysis of the anomaly's characteristics performed by the retrospective analysis module 503 can include plotting the retrospective trend of anomaly development, calculating the rate and monotonicity of anomaly development, calculating the magnitude of the deviation of the anomaly's characteristic values from the trend, and so on. In addition, the results generated by the retrospective analysis module 503 can be sent to the operator 330 of the CPS via the feedback interface 320 for analysis of the dynamics of past abnormal developments and for analysis of the causes of abnormal developments in this category.
[0136] The data generated by the retrospective analysis module 503, namely the values of the characteristics of anomalies within a specific historical time interval, can be used as supplementary input data for the predictive analysis module 504. In one aspect, the predictive analysis module 504 can be configured to predict the development of anomalies associated with individual devices and individual categories. For example, the predictive analysis module 504 can predict the values of characteristics of future anomalies. The predictive analysis module 504 can use machine learning models for predictive analysis to analyze vector data of anomaly characteristics acquired for a given historical interval. For example, the predictive analysis module 504 can use regression analysis models and extrapolation models, thereby generating predicted values of the vector of characteristics within a given time interval. Furthermore, the predictive analysis module 504 can calculate the time when certain characteristics reach a predetermined level (a critical level for anomalies of a given category). The CPS operator 330 can use the critical level of anomalies to plan maintenance and repair work.
[0137] The flow analysis module 505 can be configured to perform flow analysis on the values of the characteristics of each anomaly category. In other words, the flow analysis module 505 can be configured to analyze the values obtained at the current moment or within a given input window (input time interval). The flow analysis module 305 can record information containing the anomaly category and the values of the characteristics of each anomaly category in the anomaly database 310 in real time mode. In other words, the flow analysis module 305 can record generated information about anomalies that have occurred as information arrives. Furthermore, the flow analysis module 305 can perform a comparison of the value of the characteristic of each anomaly category with a predetermined threshold, mark an indication that the threshold has been exceeded in the anomaly database 310, and additionally send a notification indicating that the threshold has been exceeded to the CPS operator 330 via the feedback interface 320.
[0138] On the other hand, module 506, which suggests exception handling, can be configured to compare the results obtained during the analysis of the type described above performed by one of modules 503-505 with critical rules. If at least one of the critical rules is satisfied, module 506 can generate a list of actions for handling the exception based on the satisfied rules.
[0139] In certain contexts, the conditions for a critical rule may include the value of one or more CPS characteristics exceeding a given threshold, or the value of one or more CPS characteristics showing an increasing trend.
[0140] In certain aspects, the list of actions for exception handling may include, but is not limited to, the actions listed below:
[0141] a) Adjustments can be made to the data transmitters, actuators, or PID controllers that are identified as abnormal. These adjustments can be made based on CPS characteristics, according to the CPS documentation.
[0142] b) Disconnect the data transmitter, actuator, or PID controller that was identified as abnormal. For example, if the identified abnormality indicates a faulty data transmitter or that a hacker is using the data transmitter.
[0143] c) Change the computer security settings of CPS. For example, you can update various security protocols of CPS, perform a full antivirus scan, check for vulnerabilities, and disconnect vulnerable network connections.
[0144] d) Automatic correction of the control process. In one aspect, the method for corrective control can be specified based on the anomaly category.
[0145] e) Notify the SCADA system 110b' of the categories of the identified anomalies, and the results of diagnosing and monitoring each category of anomalies.
[0146] The following are some examples of implementations of the present invention.
[0147] Therefore, an anomaly category generated by the classifier module 304 during classification can include all anomalies related to the transient "stickiness" of level gauge sensors used in the viscous paraffin media characteristics of the petrochemical industry. In a given example, sensor "stickiness" means that the sensor readings periodically show zero or are generally incorrect, which is a non-critical anomaly. All these non-critical anomalies can be combined into a single category by the classifier module 304, and further analysis of this category can include diagnosing the anomalies of that category, i.e., determining the values of the characteristics of the anomalies in that category as the frequency and duration of the "stickiness," and subsequently monitoring the values of these characteristics of the anomalies in that category. Furthermore, if there are significant changes in the values of these characteristics, the corresponding information can be stored in the anomaly database 310 by the classifier module 304 and can also be sent to the CPS operator 330 via the feedback interface 320.
[0148] Another example is a situation where production technology allows at least one variable of the CPS to briefly exceed a given range of variation, but not for extended periods or excessively frequent exceedances. If such extended or excessively frequent exceedances occur, an anomaly is identified, which can be detected by one of the anomaly identification modules 301. The classifier module 304 then combines all these identified anomalies into a single category and performs subsequent diagnosis for each anomaly category by calculating values for a characteristic of the anomaly in each category. This characteristic may include, for example, the critical level of the anomaly, the frequency of anomaly occurrence, or the periodicity of anomaly occurrence. The diagnosis and monitoring module 307 can be configured to perform subsequent monitoring of the calculated values of the characteristic of the anomaly in that category (e.g., the frequency of anomaly occurrence). If the value of the characteristic "frequency of anomaly occurrence" increases over time, the diagnosis and monitoring module 307 can notify the CPS operator 330 of a critical rise in the value of that characteristic via feedback interface 320. The predictive analysis module 504 can determine this rise in the value of the characteristic "frequency of anomaly occurrence" by predicting its value, and can notify the CPS operator 330 if the prediction exceeds a given threshold for the characteristic.
[0149] Another example is internal pipe diagnostics performed using internal pipe inspection tools. Pipe walls have a wide range of defect categories, including cracks, corrosion, and dents. Diagnostic data on these defects can be used to determine values for specific characteristics selected for anomaly categories, such as the length, width, and depth of the defect. Monitoring the values of these characteristics for anomalies in that category allows pipe operators to plan maintenance and repair work, avoiding costly downtime and accidents.
[0150] Figure 6This is a flowchart illustrating an example method for diagnosing and monitoring anomalies in a CPS. In step 601, the anomaly determination module 301 identifies anomalies in the CPS 200 by analyzing the values of CPS variables. Next, in step 602, information about the anomalies found in the CPS 200 is obtained. This information may include a list of CPS variables, the values of the CPS variables over a given time interval, and the aforementioned additional information. The aggregation module 302 stores this information for each anomaly in the anomaly database 310. Subsequently, in step 603, classification features are generated for the identified anomalies based on the information collected by the generation module 302. In particular, for a specific type of CPS 200 and the processes occurring within that type of CPS 200, this set of classification features and their numerical calculation techniques can be known in advance. For example, in the case of diagnosing and monitoring anomalies in the walls of an oil pipeline using magnetic flaw detection, the size of the echo from the defect in the diagnostic data, the maximum value of the echo, the shape of the echo signal, etc., can be classification features. Then, in step 604, the classifier module 304 can classify the anomaly into at least two categories based on the classification features thus generated. Furthermore, the classification module 306 can use a set of two or more clustering or classification models to make a decision through voting among the models in that set. In step 605, the diagnosis module 501 can perform a diagnosis on each anomaly category by calculating the value of the characteristic for each anomaly category. The filtering module 502 can be configured to create rules for filtering anomalies of a category based on the diagnosis results of the diagnosis module 501. Therefore, in step 606, the diagnosis and monitoring module 307 can perform anomaly monitoring based on the diagnosis results for each anomaly category (i.e., the calculated values based on the anomaly characteristics). The above is combined... Figure 5 The diagnostics and monitoring module 307 was discussed in more detail. Previously in Figures 2-5 The specific aspects presented in it also apply to Figure 6 The method is illustrated. It should also be noted that the proposed method can also work in streaming mode, meaning that when a new exception occurs in the CPS, it can be added to an existing category or a new category can be generated for it.
[0151] Therefore, the proposed disclosure solves the aforementioned technical problem and achieves the aforementioned technical effect. Specifically, the technical effect of ensuring automatic diagnosis and monitoring of anomalies in CPS can be achieved through anomaly classification, diagnosis of anomalies in each category, and subsequent monitoring of anomalies within each anomaly category.
[0152] Figure 7 An example of a computer system on which aspects of the systems and methods disclosed herein can be implemented is shown. Computer system 20 can represent... Figure 3The system is used to diagnose and monitor anomalies and can take the form of multiple computing devices or a single computing device, such as desktop computers, laptops, handheld computers, mobile computing devices, smartphones, tablets, servers, mainframes, embedded devices, and other forms of computing devices.
[0153] As shown in the figure, computer system 20 includes a Central Processing Unit (CPU) 21, system memory 22, and a system bus 23 connecting various system components, including memory associated with the CPU 21. The system bus 23 may include bus memory or a bus memory controller, peripheral buses, and local buses capable of interacting with any other bus architecture. Examples of buses may include PCI, ISA, PCI-Express, HyperTransport, etc. TM (HyperTransport TM Unlimited bandwidth TM (InfiniBand TM Serial ATA, I2C, and other suitable interconnects. The central processing unit 21 (also called a processor) may include one or more processors with single or multiple cores. The processor 21 may execute one or more computer-executable codes that implement the techniques of this invention. The system memory 22 may be any memory used to store data used herein and / or computer programs executable by the processor 21. The system memory 22 may include volatile memory (such as Random Access Memory (RAM) 25) and non-volatile memory (such as Read-Only Memory (ROM) 24, flash memory, etc.) or any combination thereof. The Basic Input / Output System (BIOS) 26 may store basic programs for transferring information between elements of the computer system 20, such as those basic programs used when loading an operating system using ROM 24.
[0154] Computer system 20 may include one or more storage devices, such as one or more removable storage devices 27, one or more non-removable storage devices 28, or combinations thereof. The one or more removable storage devices 27 and the one or more non-removable storage devices 28 are connected to system bus 23 via storage device interface 32. In one aspect, the storage devices and corresponding computer-readable storage media are power-independent modules for storing computer instructions, data structures, program modules, and other data of computer system 20. System memory 22, removable storage devices 27, and non-removable storage devices 28 may use a wide variety of computer-readable storage media. Examples of computer-readable storage media include: machine memory such as cache, SRAM, DRAM, zero-capacitance RAM, dual-transistor RAM, eDRAM, EDO RAM, DDR RAM, EEPROM, NRAM, RRAM, SONOS, PRAM; flash memory or other storage technologies such as in solid-state drives (SSDs) or flash drives; magnetic tape cassettes, magnetic tapes, and disk storage such as in hard disk drives or floppy disks; optical storage such as in optical discs (CD-ROMs) or digital versatile optical discs (DVDs); and any other media that can be used to store desired data and can be accessed by the computer system 20.
[0155] The system memory 22, removable storage device 27, and non-removable storage device 28 of computer system 20 can be used to store the operating system 35, additional application programs 37, other program modules 38, and program data 39. Computer system 20 may include a peripheral interface 46 for transmitting data from input devices 40, such as a keyboard, mouse, stylus, game controller, voice input device, touch input device, or other peripheral devices, such as a printer or scanner via one or more I / O ports, such as serial ports, parallel ports, Universal Serial Bus (USB), or other peripheral interfaces. Display devices 47 (such as one or more monitors, projectors, or integrated displays) can also be connected to system bus 23 via output interface 48 (such as a video adapter). In addition to display devices 47, computer system 20 may also be equipped with other peripheral output devices (not shown), such as speakers and other audiovisual equipment.
[0156] Computer system 20 can operate in a networked environment using a network connection to one or more remote computers 49. The one or more remote computers 49 can be local computer workstations or servers, including most or all of the elements described above in the description of the nature of computer system 20. Other devices may also be present in the computer network, such as, but not limited to, routers, network sites, peer devices, or other network nodes. Computer system 20 may include one or more network interfaces 51 or network adapters for communicating with remote computers 49 via one or more networks, such as a local area network (LAN) 50, a wide area network (WAN), an intranet, and the Internet. Examples of network interfaces 51 may include Ethernet interfaces, Frame Relay interfaces, SONET (Synchronous Fiber Network) interfaces, and wireless interfaces.
[0157] Various aspects of the present invention can be systems, methods, and / or computer program products. A computer program product may include one or more computer-readable storage media having computer-readable program instructions on which a processor performs various aspects of the present invention.
[0158] Computer-readable storage media can be tangible devices that hold and store program code in the form of instructions or data structures, which can be accessed by a processor of a computing device (such as computer system 20). Computer-readable storage media can be electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. As examples, such computer-readable storage media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), portable optical disc read-only memory (CD-ROM), digital versatile optical disc (DVD), flash memory, hard disk, portable computer disk, memory stick, floppy disk, or even mechanical encoding devices, such as punch cards or raised structures in recesses on which instructions are recorded. As used herein, computer-readable storage media should not be considered as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or transmission media, or electrical signals transmitted through wires.
[0159] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing device, or downloaded via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. This network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network interface in each computing device receives and forwards the computer-readable program instructions from the network for storage in a computer-readable storage medium within the respective computing device.
[0160] Computer-readable program instructions used to perform the operations of this invention can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages (including object-oriented programming languages and traditional procedural programming languages). The computer-readable program instructions (as a standalone software package) can be executed entirely on the user's computer, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network (including LAN or WAN) or can be connected to an external computer (e.g., via the Internet). In some embodiments, electronic circuits (including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs)) can execute the computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuits and thereby perform aspects of the invention.
[0161] In various respects, the systems and methods described in this invention can be processed in modules. As used herein, the term "module" refers to, for example, a real-world device, component, or arrangement of components implemented using hardware (e.g., by an application-specific integrated circuit (ASIC) or FPGA), or a combination of hardware and software, such as a combination implemented by a microprocessor system and an instruction set that implements the module's functionality (which, when executed, translates the microprocessor system into a special-purpose device). A module can also be implemented as a combination of two modules, wherein certain functions are facilitated solely by hardware, and other functions are facilitated by a combination of hardware and software. In some implementations, at least a portion of the module (and in some cases, all of the module) can run on a processor of a computer system. Therefore, each module can be implemented in a variety of suitable configurations and should not be limited to any particular implementation illustrated herein.
[0162] For clarity, not all routine features of each aspect are disclosed herein. It should be understood that in the development of any practical implementation of this invention, many specific implementation decisions must be made to achieve the developer's specific objectives, and these specific objectives will vary for different implementations and different developers. It should be understood that such development efforts can be complex and time-consuming, but remain a routine engineering task for those skilled in the art who understand the advantages of this invention.
[0163] Furthermore, it should be understood that the wording or terminology used herein is for descriptive rather than limiting purposes, and therefore the terminology or terminology in this specification should be interpreted by a person skilled in the art based on the teachings and guidance set forth herein in conjunction with the knowledge of one or more persons skilled in the art. Moreover, it is not intended that any term in this specification or claims be assigned an uncommon or specific meaning unless expressly stated so.
[0164] The various aspects disclosed herein include present and future known equivalents of the known modules cited herein in an illustrative manner. Furthermore, although various aspects and applications have been shown and described, it will be apparent to those skilled in the art, who will appreciate the advantages of the invention, that further modifications are possible beyond what has been mentioned above without departing from the inventive concept disclosed herein.
Claims
1. A method for diagnosing and monitoring anomalies in a cyber-physical system (CPS), the method comprising: obtaining information related to an anomaly identified in the CPS, wherein the obtained information comprises at least one value of one or more CPS variables; generating one or more classification features of the anomaly identified in the CPS based on the obtained information; classifying the anomaly identified in the CPS into two or more anomaly classes based on the generated classification features, wherein each of the two or more anomaly classes is associated with one or more anomaly characteristics; performing diagnosis of the anomaly in each of the two or more anomaly classes by computing values of the anomaly characteristics associated with each of the two or more anomaly classes; and monitoring the anomaly in each of the two or more anomaly classes by performing at least one of the following types of analysis using the computed values of the anomaly characteristics associated with each of the two or more anomaly classes at a predetermined frequency: retrospective analysis of values of additional characteristics of the two or more anomaly classes over a predetermined historical time period; predictive analysis of values of the anomaly characteristics associated with each of the two or more anomaly classes; and stream analysis of values of the anomaly characteristics associated with each of the two or more anomaly classes.
2. The method of claim 1, wherein, identifying the anomaly by performing the following steps: predicting at least one value of the one or more CPS variables; determining a total prediction error based on the predicted at least one value of the one or more CPS variables; and identifying the anomaly if the determined total prediction error exceeds a predetermined threshold. identifying the anomaly by performing the following steps:
3. The method of claim 1, wherein, identifying the anomaly by applying a trained machine learning model to the at least one value of the one or more CPS variables. identifying the anomaly by performing the following steps:
4. The method of claim 1, wherein, determining whether at least one value of the one or more CPS variables is outside the bounds of a previously specified value range of a corresponding CPS variable; and identifying the anomaly in response to determining that the value of at least one of the one or more CPS variables is outside the bounds of the previously specified value range of the corresponding CPS variable. the obtained information further comprises at least one of: a time interval at which the detected anomaly was observed, a contribution of each of the one or more CPS variables to the detected anomaly, information about a detection method of the detected anomaly, at least one value of the one or more CPS variables at each time instant of the observation time interval. for each of the one or more CPS variables, the obtained information further comprises: a time series of values of the corresponding CPS variable; a current magnitude of deviation of a predicted CPS variable value from an actual CPS variable value; a smoothed value of deviation of the predicted CPS variable value from the actual CPS variable value.
5. The method of claim 1, wherein, 6. The method of claim 1, wherein, 7. The method of claim 1, wherein, The at least one value of the one or more CPS variables comprises at least one of: a measured value of a data transmitter; a value of a manipulated variable of an actuator; one or more values of an input signal of a proportional-integral-derivative, PID, controller; a value of an output signal of a PID controller. The at least one value of the one or more CPS variables comprises at least one of: a measured value of a data transmitter; a value of a manipulated variable of an actuator; one or more values of an input signal of a proportional-integral-derivative, PID, controller; a value of an output signal of a PID controller.
8. The method of claim 1, wherein, Monitoring the CPS to detect anomalies further comprises comparing results of the performed analysis to one or more critical rules, and in response to determining that at least one of the critical rules is satisfied, generating a list of actions for handling the detected anomaly according to the satisfied critical rule.
9. The method of claim 1, wherein, The one or more classification features are generated by assigning the one or more CPS variables to each of the one or more classification features.
10. The method of claim 1, wherein, The one or more anomaly characteristics comprise at least one of: a calculated value of a critical level of a particular anomaly, a frequency of occurrence of a particular class of anomalies, a periodicity of occurrence of a particular class of anomalies.
11. The method of claim 1, wherein, The one or more classification features are generated based on feedback from an operator of the CPS.
12. The method of claim 1, wherein, Performing classification of the identified anomalies further comprises performing the classification using one of a trained classification model and a trained clustering model, wherein input data for the trained classification model or the trained clustering model comprises the one or more classification features, and wherein results of the classification comprise assigning an anomaly class to each of the identified anomalies.
13. A system for diagnosing and monitoring anomalies in a cyber-physical system, CPS, the system comprising: a memory and a hardware processor configured to: obtain information related to an anomaly identified in the CPS, wherein the obtained information comprises at least one value of one or more CPS variables; generate one or more classification features of the anomaly identified in the CPS based on the obtained information; classify the anomaly identified in the CPS into two or more anomaly classes based on the generated classification features, wherein each of the two or more anomaly classes is associated with one or more anomaly characteristics; perform diagnosis of anomalies in each of the two or more anomaly classes by calculating values of the anomaly characteristics associated with each of the two or more anomaly classes; and monitor anomalies in each of the two or more anomaly classes by performing at least one of the following types of analysis using the calculated values of the anomaly characteristics associated with each of the two or more anomaly classes at a predetermined frequency: a retrospective analysis of values of additional characteristics of the two or more anomaly classes over a predetermined historical time period; a predictive analysis of values of the anomaly characteristics associated with each of the two or more anomaly classes; and a streaming analysis of values of the anomaly characteristics associated with each of the two or more anomaly classes.
14. The system of claim 13, wherein, the hardware processor configured to identify anomalies is further configured to: predicting at least one value of the one or more CPS variables; determining a total prediction error based on the predicted at least one value of the one or more CPS variables; and identifying an anomaly if the determined total prediction error exceeds a predetermined threshold.
15. The system of claim 13, wherein, The hardware processor configured to identify an anomaly is further configured to: identify an anomaly by applying a trained machine learning model to the at least one value of the one or more CPS variables.
16. The system of claim 13, wherein, The hardware processor configured to monitor the CPS to identify an anomaly is further configured to: determine whether at least one value of the one or more CPS variables is outside the limits of a previously specified value range of a respective CPS variable; and identify an anomaly in response to determining that a value of at least one of the one or more CPS variables is outside the limits of a previously specified value range of the respective CPS variable.
17. The system of claim 13, wherein, The obtained information further comprises at least one of: a time interval in which the detected anomaly is observed, a contribution of each of the one or more CPS variables to the detected anomaly, information about a detection method of the detected anomaly, at least one value of the one or more CPS variables at each time instant of the observation time interval.
18. The system of claim 13, wherein, For each of the one or more CPS variables, the obtained information further comprises: a time series of respective CPS variable values; a current size of a deviation of a predicted CPS variable value from an actual CPS variable value; a smoothed value of the deviation of the predicted CPS variable value from the actual CPS variable value.
19. A non-transitory computer-readable medium having stored thereon computer- executable instructions for diagnosing and monitoring anomalies in a cyber-physical system (CPS), the computer-executable instructions comprising instructions for: obtaining information related to the anomaly identified in the CPS, wherein, obtaining at least one value of one or more CPS variables; generating one or more classification features of an anomaly identified in the CPS based on the obtained information; classifying the anomaly identified in the CPS into two or more anomaly classes based on the generated classification features, wherein each of the two or more anomaly classes is associated with one or more anomaly characteristics; performing a diagnosis of an anomaly in each of the two or more anomaly classes by calculating a value of an anomaly characteristic associated with each of the two or more anomaly classes; and monitoring the anomaly in each of the two or more anomaly classes by performing at least one type of analysis from the following types of analysis at a predetermined frequency using the calculated value of the anomaly characteristic associated with each of the two or more anomaly classes: a retrospective analysis of values of additional characteristics of the two or more anomaly classes over a predetermined historical time period; a predictive analysis of values of the anomaly characteristic associated with each of the two or more anomaly classes; and a predictive analysis of values of the anomaly characteristic associated with each of the two or more anomaly classes. stream analysis of values of an anomaly characteristic associated with each of the two or more anomaly categories.
Citation Information
Patent Citations
Anomaly Detection for Cyber-Physical Systems
US20210089661A1