Virtual DCS safety operator for accident detection and response
By monitoring and jointly analyzing IT and OT data in a cloud-on-premises DCS, and applying specific rules to a virtual DCS security operator, the problem of rapid and automated security incident detection and response is solved, improving security and detection efficiency.
Patent Information
- Application Number
- CN202510563075.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-03
- Filing Date
- 2025-04-30
- Publication Date
- 2025-11-04
AI Technical Summary
In cloud-local distributed control systems (DCS), there are challenges in data processing and information isolation for security incident detection and response, which need to be improved to achieve rapid and automated security incident response.
By monitoring information technology (IT) and operational technology (OT) data during the production process, jointly analyzing and applying domain-specific security incident detection and response rules, and utilizing a virtual DCS security operator to work autonomously in a containerized environment, rapid detection and response to security incidents can be achieved.
It improves the reliability and speed of security incident detection, reduces false alarms and missed alarms, achieves seamless integration with cloud-local DCS, and reduces the user burden through autonomous operation system.
Smart Images

Figure CN120893038A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to a virtual DCS security operator for incident detection and response. BACKGROUND
[0002] Safety incident detection in cloud-on-premise distributed control systems (DCS) is challenging as a large amount of data has to be processed and there is a need for segregated handling of information technology (IT) related data (e.g. like user access logs) and operational technology (OT) related data (e.g. like motor start irregularities). Upon detection of a potential incident, a security management system has to react quickly to possibly contain the incident and prevent it from further spreading through the system.
[0003] Therefore, there are several possible drawbacks with respect to safety incident detection in cloud-on-premise DCS. Therefore, there is room for improvement and a need. In particular, there is a need for an automated safety incident response that is able to contain safety breaches. SUMMARY
[0004] In view of the above, it is an object of the present disclosure to overcome at least part of the possible drawbacks in safety incident detection in cloud-on-premise DCS.
[0005] A system that overcomes at least part of these drawbacks can need to cover specific requirements with respect to data handling and functionality. For example, it can be needed that the system is able to handle and correlate IT related and OT related data. Unlike general intrusion detection or incident monitoring systems, for a DCS there is a need for domain specific incident detection and response rules. The system can need to work in a cloud-on-premise environment to leverage containerized DCS services and needs to work mainly autonomously to not overburden the user. According to several examples, the virtual DCS security operator disclosed in the present application can cover these requirements.
[0006] Thus, in view of the above problems and in order to solve one or more disadvantages, in a first aspect, a method for security incident detection in a cloud-on-premise distributed control system, DCS, in an industrial process automation is provided. The method comprises monitoring information technology, IT, related data and operation technology, OT, related data at a production process and a containerized DCS associated with the production process. The method further comprises jointly analyzing first data and second data, the first data being indicative of a first monitored data from the monitoring of the IT related data, the second data being indicative of a second monitored data from the monitoring of the OT related data. The joint analysis is based on associating at least part of the first data with at least part of the second data and / or based on associating at least part of the second data with at least part of the first data. The method further comprises, based on the joint analysis, detecting a security incident in consideration of predetermined security incident detection rules. The method further comprises, based on a result of the detecting, responding to the detected security incident for handling the detected security incident in consideration of predetermined security incident response rules.
[0007] As an example for improving understandability, the correlation can comprise correlating an OT event, such as a sensor value drift, with an IT event, such as a user login, by a timestamp, so that a potential manipulation of the automation process can be detected.
[0008] It should be noted that the joint analysis can also be understood as a combined analysis. The joint analysis can also be understood as a separate or sequential analysis of the first data and the second data, but a comparative or cross analysis of the results from the separate or sequential analysis. That is, the first data is not analyzed separately or exclusively, and the second data is not analyzed separately or exclusively.
[0009] It should also be noted that the predetermined security incident detection rules can be customized security incident detection rules and can be specified for a particular DCS and / or production process. The expression “responding for handling” can be understood as taking or triggering measures to handle, e.g. to eliminate, the detected security incident in response to detecting the security incident. These measures can comprise, for example, a notification can be issued to a user or the detected security incident is autonomously eliminated by the data processing device or data processing system.
[0010] The method according to the first aspect is advantageous in that it can contribute to a higher security by a better security detection. A potential fast reaction to security breaches is also achieved by the automatic reaction. Furthermore, the user interaction and improvement are improved since an autonomous running system with user interaction and improvement capabilities is provided. Moreover, a seamless integration with a cloud-on-premise DCS is achieved.
[0011] According to some examples of the present disclosure, the method can further comprise performing the monitoring, the joint analysis, the detecting, and the responding by a virtual DCS security operator. According to some examples of the present disclosure, the virtual DCS security operator can be a software agent running in the container orchestration cluster; and / or the virtual DCS security operator can be an autonomously running security operator.
[0012] Thus, due to the constantly running software agent, the user can be liberated and more security incidents can be detected, security incidents can be detected more reliably, and security incidents can be eliminated more quickly. Thus, the overall security is improved.
[0013] According to some examples of the present disclosure, the first monitoring data can represent monitored IT-related data comprising monitored system diagnostic data, and wherein the second monitoring data can represent monitored OT-related data comprising monitored process data. Additionally or alternatively, the correlating can comprise correlating at least part of the first monitoring data with at least part of the second monitoring data and / or correlating at least part of the second monitoring data with at least part of the first monitoring data.
[0014] Thus, due to the knowledge gain from correlating the monitored IT data and OT data, the security incident detection can be more comprehensive and can provide more insights.
[0015] According to some examples of the present disclosure, the monitoring of the IT-related data and the OT-related data can comprise monitoring the IT-related data and the OT-related data taking into account predetermined security incident monitoring rules. Additionally or alternatively, the monitoring of the IT-related data and the OT-related data can comprise monitoring data from a kubernets application programming interface, API, server and from an open platform communication unified architecture, OPC UA, server. Additionally or alternatively, the method can further comprise accessing the Kubernetes API server and performing the responding based on adjusting parameters available in the Kubernetes API server.
[0016] Thus, a large and varying amount of data can be used and acquired as a basis for the monitoring. Thus, the quality and reliability of the security incident monitoring and detection is further increased. This includes increasing the amount of true positives (i.e. actual security incidents that are correctly detected) and decreasing the amount of false positives (i.e. events that are incorrectly marked as security incidents by the detection system) and false negatives (i.e. attacks that are missed by the detection system). Thus, the responding to the detected security incidents is further increased as more options for handling the detected security incidents can be considered.
[0017] Further, according to several examples of the present disclosure, based on accessing the Kubernetes API server, the response can comprise reconfiguring any Kubernetes resources, adding and / or removing nodes from the cluster associated with the DCS, starting and / or stopping pods in these nodes, and / or changing network routing between components in the cluster.
[0018] Further, according to several examples of the present disclosure, the method can further comprise interacting with the production process and interacting with IT, wherein at least one of monitoring, detecting, and responding can be based on the interaction.
[0019] It should be noted that the interaction with the production process can comprise invoking an OPC UA method or writing a setpoint to an OPC UA server. The interaction with IT can comprise detaching a node from the cluster or forcefully killing a compromised supervisory component via Kubernetes.
[0020] Thus, the quality of the security incident monitoring and detection and response is further improved.
[0021] According to several examples of the present disclosure, the joint analysis can comprise detecting a security incident in one of the first data and the second data; and analyzing the other of the first data and the second data for an event associated with the detected security incident.
[0022] Thus, the efficiency of the data analysis is improved, because if an indication of a security incident is given in one of the first data and the second data, a cross analysis between the first and the second data is performed. Thus, continuous cross analysis can be avoided.
[0023] According to several examples of the present disclosure, the production process and the containerized DCS can correspond to a specific domain, and wherein the predetermined security incident detection rules and the predetermined security incident response rules can be specific to the specific domain.
[0024] Thus, the predetermined security incident detection rules can be understood as specified or specific predetermined security incident detection rules. The predetermined security incident response rules can be understood as specified or specific predetermined security incident response rules. For example, DCS-specific security detection and resolution rules encoded in custom operation extensions, e.g. Kubernetes custom resource definitions, can be applied.
[0025] Thus, more suitable, more individual, and more applicable rules can be applied. Thus, the quality and reliability of the security incident monitoring and detection and response is further increased.
[0026] According to some examples of the present disclosure, the method can further comprise using a virtual DCS security custom resource comprising at least part of predetermined security incident detection rules, at least part of predetermined security incident response rules, and at least part of predetermined security incident monitoring rules.
[0027] For example, the virtual DCS security custom resource can be a dedicated ConfigMap, e.g. like key-value pairs holding DCS security incident monitoring rules, security incident detection rules, and security incident response rules. According to some examples of the present disclosure, a rule can declare that more than five individual logins to an OPC UA server per day are unusual and can indicate a security incident. A security incident response rule can declare to disconnect all clients from the OPC UA server and to temporarily forbid further user access. Other security incident response rules can be to change passwords or to reconfigure components.
[0028] Thus, based on the virtual DCS security custom resource, dynamic configuration can be provided by custom operations, e.g. like Kubernetes custom resources, which allow to improve incident detection and resolution over the system lifecycle.
[0029] According to some examples of the present disclosure, the method can further comprise modifying the virtual DCS security custom resource for at least one of the predetermined security incident detection rules, the predetermined security incident response rules, and the predetermined security incident monitoring rules. The modification can be performed manually by a user, automatically by an inference system associated with the DCS without involving the user, or semi-automatically with the user providing guidance to the inference system.
[0030] It is noted that the automation of the automatic or semi-automatic modification can be implemented based on applying a machine learning / inference system to existing data to derive new rules. The existing data can comprise the first data and the second data as indicated above. The existing data can further comprise historical first data and historical second data. The existing data can further comprise externally obtained data, e.g. from a Security Information and Event Management, SIEM, system. The existing data can further comprise predetermined security incident detection, response, and monitoring rules. The new rules can comprise new security incident detection, response, and monitoring rules.
[0031] Thus, the security incident monitoring and detection, and the response are further individualized, and the user is further relieved, as the modification can be performed automatically or semi-automatically. Thus, the quality of the security incident resolution is improved.
[0032] According to some examples of the present disclosure, the response can comprise at least one of:
[0033] - informing a user about the detected security incident;
[0034] - The application receives commands from the user regarding detected security incidents;
[0035] -Autonomously apply information from pre-defined security incident response rules regarding detected security incidents.
[0036] Therefore, the pre-determined safety incident response rules; and
[0037] - Simulate the response to the detected security incident before executing the response, and then execute the response based on the simulation results.
[0038] For example, a response could include redeploying potentially compromised Kubernetes pods to a secure isolation zone within the cluster. In more severe cases, a response could include partially shutting down all non-security-critical pods in the system, such as oversight DCS pods, to accommodate a security incident. Based on several examples in this disclosure, predefined security incident response rules can identify non-security-critical pods.
[0039] For example, certain incident responses can be simulated before they are executed. Simulations can include replicas of DCS pods or even virtual pods, and can help evaluate the consequences of a partial system shutdown.
[0040] Therefore, the quality of the response was further improved.
[0041] According to several examples of this disclosure, the method may further include exchanging third data with a security information and event management system (SIEM); and further performing at least one of monitoring, joint analysis, detection, and response based on the third data. The third data indicates at least one of the following:
[0042] - Recorded events that occur during production and / or at containerized DCS.
[0043] - The response executed
[0044] -Additional, removed, and / or updated security incident monitoring rules,
[0045] - Additional, removed, and / or updated security incident detection rules, and
[0046] - Additional, removed, and / or updated security incident response rules.
[0047] Therefore, further insights have been gained due to the revised rules for monitoring, detecting, and responding to safety incidents, and the handling of safety incidents has been further improved, potentially becoming more specific and individualized.
[0048] According to a second aspect, a data processing device for security incident detection in a cloud-on-premise DCS in industrial process automation is provided. The data processing device comprises a processor configured to perform the method of the first aspect.
[0049] An advantage of the data processing device according to the second aspect is that it can contribute to a higher security by better security detection. A potentially fast reaction to security breaches is also achieved by the automatic reaction. Furthermore, the user interaction and improvement are improved since an autonomous running system with user interaction and improvement capabilities is provided. Moreover, a seamless integration with the cloud-on-premise DCS is achieved.
[0050] According to several examples of the present disclosure, the data processing device can comprise a Kubernetes client, an OPC UA client, an incident detector, an incident responder, and a user interface. The Kubernetes client can be in communicative connection with the incident detector and the incident responder, the OPC UA client can be in communicative connection with the incident detector and the incident responder, the incident detector can be in communicative connection with the user interface and the incident responder, and the incident responder can be in communicative connection with the user interface and the incident detector. The Kubernetes client and the OPC UA client can be configured to perform the monitoring according to the method of the first aspect, the incident detector can be configured to perform the joint analysis and detection according to the method of the first aspect, and the incident responder can be configured to perform the response according to the method of the first aspect.
[0051] According to a third aspect, a security system or a data processing system for security incident detection in a cloud-on-premise DCS in industrial process automation is provided. The data processing system comprises the data processing device of the second aspect configured to perform the method of the first aspect. Additionally or alternatively, the data processing system comprises means for performing the method of the first aspect.
[0052] An advantage of the data processing system according to the third aspect is that it can contribute to a higher security by better security detection. A potentially fast reaction to security breaches is also achieved by the automatic reaction. Furthermore, the user interaction and improvement are improved since an autonomous running system with user interaction and improvement capabilities is provided. Moreover, a seamless integration with the cloud-on-premise DCS is achieved.
[0053] According to a fourth aspect, an industrial plant comprising the data processing device of the second aspect and / or the data processing system of the third aspect is provided, the data processing device being configured to perform the method of the first aspect.
[0054] According to several examples, an “industrial plant” can refer to an industrial plant or industrial production plant comprising one or more pipelines, production lines and / or assembly lines for converting one or more educts into a product and / or for assembling one or more components into an end product. According to several examples, it can refer to an industrial plant in the oil industry, the natural gas industry or the chemical industry.
[0055] According to the fourth aspect, the industrial plant can participate in achieving a higher security by better security detection. A potentially fast reaction to security breaches is also achieved by the automatic reaction. Furthermore, the user interaction and improvement are improved since an autonomous running system with user interaction and improvement capabilities is provided. Moreover, a seamless integration with a cloud-on-premise DCS is achieved.
[0056] According to the fifth aspect, a computer-readable medium comprising instructions which, when executed by a computing system, cause the computing system to perform the method of the first aspect is provided. The computer-readable medium can be transitory or non-transitory, volatile or non-volatile.
[0057] According to the fifth aspect, the computer-readable medium can participate in achieving a higher security by better security detection. A potentially fast reaction to security breaches is also achieved by the automatic reaction. Furthermore, the user interaction and improvement are improved since an autonomous running system with user interaction and improvement capabilities is provided. Moreover, a seamless integration with a cloud-on-premise DCS is achieved.
[0058] According to the sixth aspect, a computer program product comprising instructions which, when executed by a computing system, enable or cause the computing system to perform the method of the first aspect is provided. The computer program product can comprise a computer-readable medium comprising the instructions of the computer program product.
[0059] According to the sixth aspect, the computer program product can participate in achieving a higher security by better security detection. A potentially fast reaction to security breaches is also achieved by the automatic reaction. Furthermore, the user interaction and improvement are improved since an autonomous running system with user interaction and improvement capabilities is provided. Moreover, a seamless integration with a cloud-on-premise DCS is achieved.
[0060] According to the seventh aspect, a use of the data processing device of the second aspect, and / or the data processing system of the third aspect, and / or the industrial plant of the fourth aspect is provided.
[0061] The use according to the seventh aspect is advantageous in that it can contribute to a higher security by better security detection. Also a potentially fast reaction to security breaches is achieved by the automatic reaction. Further, the user interaction and improvement is improved since an autonomous running system with user interaction and improvement capabilities is provided. Further, a seamless integration with cloud-on-premise DCS is achieved.
[0062] The method of the first aspect can be computer implemented.
[0063] Optional features of the first aspect can form part of any of the second to seventh aspects mutatis mutandis.
[0064] As used herein, the term “obtaining” can include, for example, receiving from another system, device, or process; receiving via an interaction with a user; loading or retrieving from storage or memory; measuring or capturing using a sensor or other data acquisition device.
[0065] As used herein, the term “determining” encompasses a wide variety of actions and can include, for example, calculating, computing, processing, deriving, investigating, looking up (such as, for example, in a table, a database or another data structure), ascertaining and the like. Also, “determining” can include receiving (such as, for example, receiving information), accessing (such as, for example, accessing data in a memory) and the like. Also, “determining” can include resolving, selecting, choosing, establishing and the like.
[0066] The indefinite articles “a” or “an” do not exclude a plurality. Furthermore, a plurality is intended unless the context clearly indicates to the contrary. For example, the phrases “a first” and “a second” are not excluded from the scope of the present application. Furthermore, the terms “first” and “second” are used herein for clarity to describe different subjects for the purpose of the discussion only and are not limiting. Thus, a subject can be regarded as the first or second subject according to the context.
[0067] The phrases “one or more of the items,” “at least one of the items,” and “one, two, or more of the items” as used herein covers all of the following interpretations of the above phrases: the item is the one and only one of the items; the item is one of the items; and the item is two or more of the items.
[0068] The term “comprising”, used in the detailed description and throughout the claims, should not be interpreted as being restricted to security systems, devices, methods, features or steps that might be explicitly disclosed from the context. The term “comprising” does not exclude any other steps.
[0069] The application can include one or more aspects, examples, or features alone or in combination. Any optional feature of one of the above aspects or sub-aspects is applicable mutatis mutandis to any other aspect.
[0070] The foregoing aspects will become apparent from the following detailed description, provided with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0071] The detailed description will now be made with reference to the accompanying drawings, of which:
[0072] Figure 1 Contexts of a virtual DCS security operator are shown according to several examples of the present disclosure;
[0073] Figure 2 Flowcharts representing methods according to several examples of the present disclosure are shown;
[0074] Figure 3 Systems including a virtual DCS security operator and a SIEM according to several examples of the present disclosure are shown; and
[0075] Figure 4 Architectures of a virtual DCS security operator according to several examples of the present disclosure are shown. DETAILED DESCRIPTION
[0076] According to several examples of the present disclosure, a virtual DCS security operator for a cloud-on-premise distributed control system (DCS) for use in process automation is provided. The virtual DCS security operator is a continuously running software agent, running in a container orchestration cluster, and can continuously monitor IT data and OT data to discover potential security incidents. Thus, the virtual DCS security operator can monitor data from the production process and the DCS, can detect potential security incidents according to domain-specific rules, and in specific cases, can react to them autonomously. Thus, the virtual DCS security operator can be configured with domain-specific detection rules. In more detail, upon detecting a security breach, the virtual DCS security operator can query a user for an incident response, or autonomously execute a pre-specified, domain-specific incident response. For the autonomous reaction, the operator can use pre-specified, domain-specific rules, and can issue commands to the IT infrastructure (such as turning off a server, for example) or to the OT infrastructure (such as reconfiguring a heat exchanger, for example). Thus, the virtual DCS security operator is able to quickly react to security breaches, and can potentially prevent them from spreading. The configuration of the virtual DCS security operator can be extended during runtime without interrupting service, thus the incident detection and response can continuously become more powerful.
[0077] According to several examples of the present disclosure, the virtual DCS safety operator is a virtualized operator for safety incident monitoring, detection, and automated reaction. The virtual DCS safety operator can correlate both IT data and OT data during incident monitoring in order to be able to detect more subtle safety breaches. The virtual DCS safety operator can include a dynamic configuration by custom operations, such as e.g. Kubernetes custom resources, in order to be able to improve incident detection and resolution over the system lifecycle. The virtual DCS safety operator can leverage both automation devices, e.g. via Open Platform Communications Unified Architecture (OPC UA), and IT devices, e.g. via Kubernetes, to perform autonomous incident response.
[0078] According to several examples of the present disclosure, in more detail, in case it is deployed in a container orchestration framework, the virtual DCS safety operator can quickly react to potential safety incidents. This can help to quickly contain safety breaches in the system and isolate them. The virtual DCS safety operator can correlate process data and system diagnostic data and can therefore potentially identify a larger variety of safety breaches than other safety incident reporting systems that rely on system diagnostic data only. For example, it can correlate OT events, such as sensor value drifts, with IT events, such as user logins, by timestamps, so that potential manipulation of the automation process can be detected. It can also apply DCS-specific safety detection and resolution rules encoded in custom operation extensions, such as e.g. Kubernetes custom resource definitions. Unlike traditional reporting systems, the virtual DCS safety operator can interact with the process, e.g. call OPC UA methods or write setpoints to an OPC UA server, and can interact with IT, e.g. detach nodes from the cluster via Kubernetes, or force-kill compromised supervisory components.
[0079] Reference is now made to Figure 1 , Figure 1 A typical context or context system 100 in which the virtual DCS safety operator 121 can run is shown. The virtual DCS safety operator 121 is started at cluster startup and runs in the control plane 120 of the cluster, as a dedicated Kubernetes management node including a Kubernetes controller manager 123, a Kubernetes API server 124, and a Kubernetes scheduler 125. The control plane 120 can be hosted on multiple nodes 130, 140, 150 to provide redundancy for higher robustness, in which case the virtual DCS safety operator 121 will also feature multiple instances running in parallel, one of which is the leader and the others are followers, ready to take over in case the leader fails.
[0080] The virtual DCS security operator 121 uses a virtual DCS security customization resource 122. In this case, the virtual DCS security customization resource 122 can be a dedicated ConfigMap, such as key-value pairs storing DCS security incident monitoring rules, security incident detection rules, and security incident response rules. For example, a rule could state that more than five individual logins to the OPC UA server per day are unusual and could indicate a security incident. A response rule could state that all clients are disconnected from the OPC UA server and further user access is temporarily prohibited. Other response rules could be changing passwords or reconfiguring components. More extreme response rules could involve shutting down running pods in the event of an operating system-level security vulnerability. Figure 1 The pods shown as examples are 131, 132, 141, 142, 143, 151, 152, 153, or exhausting the entire nodes 130, 140, 150. The system has predefined, generic DCS rules that can be applied to most cloud-native DCS systems. Users 110 can add their own custom rules declaratively via standard Kubernetes tools (such as CLI or dashboards) or via dedicated configuration tools written to the Kubernetes API.
[0081] To monitor system and application responsiveness, the virtual DCS security operator 121 can access the Kubernetes API server 124. For example, the virtual DCS security operator 121 can start or stop pods 131, 132, 141, 142, 143, 151, 152, and 153, reconfigure any Kubernetes resources, add or remove nodes 130, 140, and 150 from the cluster, or change network routing between components. The Kubernetes API provides a rich interface to all types of system management functions related to cloud-native software. Changes to the API are picked up by the Kubernetes scheduler 125, which, for example, schedules the creation or deletion of pods 131, 132, 141, 142, 143, 151, 152, and 153 on one of the worker nodes 130, 140, and 150.
[0082] Dozens of worker nodes 130, 140, and 150 can execute DCS application services within software containers orchestrated as pods 132, 142, 143, 151, 152, and 153. These application services within the DCS include alarm management, process graphics, process history logging, control execution, and more. For all industrial assets managed by the system, the DCS typically includes an asset catalog (see...). Figure 1The asset catalog provides different views on assets following the object-oriented paradigm (see pod 131 in FIG. 1). Real-time data from the process (e.g. sensor values or machine states) can be fetched via references provided by the asset catalog. The virtual DCS safety operator 121 is configured to monitor selected variables of selected assets that can be potentially relevant for incident detection. Real-time data can be transmitted via typical communication protocols like OPC UA. To discover OPC UA servers, the virtual DCS safety operator 121 can also connect to an OPC UA Global Discovery Server (GDS) according to IEC 62541-12 (see pod 141 in FIG. 1). The OPC UA GDS provides a catalog of all OPC UA servers registered in the system. Figure 1
[0083] Reference is now made to Figure 2 , Figure 2 a flowchart indicating a method according to several examples of the present disclosure is shown. The method is a method for safety incident detection in a cloud-on-premise DCS in industrial process automation.
[0084] The method according to Figure 2 may be applied by the virtual DCS safety operator 121 as described above with reference to Figure 1 .
[0085] The method starts in S200.
[0086] In S210, the method comprises monitoring information technology, IT, related data and operation technology, OT, related data at a production process and a containerized DCS associated with the production process.
[0087] In S220, the method comprises jointly analyzing first data D1 and second data D2, the first data D1 being indicative of a first monitored data from the monitoring of the IT related data, the second data D2 being indicative of a second monitored data from the monitoring of the OT related data, the joint analysis being based on associating at least part of the first data D1 with at least part of the second data D2 and / or based on associating at least part of the second data D2 with at least part of the first data D1.
[0088] In S230, the method comprises detecting a safety incident based on the joint analysis in consideration of predetermined safety incident detection rules.
[0089] In S240, the method comprises responding to the detected safety incident based on the detection result to handle the detected safety incident in consideration of predetermined safety incident response rules.
[0090] The method ends in S250.
[0091] Reference is now made toFigure 3 , Figure 3 The virtual DCS safety operator 121 is schematically shown to interface with a regular security incident and event management (SIEM) system 310 that can be provided at the cloud server 300. The exchange between the two systems works in both directions, i.e. data D3 can be exchanged in both directions between the virtual DCS safety operator 121 and the SIEM 310. The virtual DCS safety operator 121 can send logged events and executed responses to the SIEM system 310 to be displayed in a corresponding user interface of the SIEM system 310 used by e.g. a network security expert. The SIEM system 310 can provide new incident detection and response rules to the virtual DCS safety operator 121 based on e.g. learning from other systems. This information can then be used by the virtual DCS safety operator 121 for fast, more informed incident detection and response.
[0092] Reference is now made to Figure 4 , Figure 4 A high-level internal structure of the virtual DCS safety operator 121 is described. The virtual DCS safety operator 121 is continuously executed in three different phases, namely a monitoring phase, a detection phase and a response phase, wherein according to several examples of the present disclosure the detection or detection phase can be understood to also comprise an analysis or analysis phase in which data obtained from the monitoring or monitoring phase is analyzed. The monitoring can be independent from the detection and response and can continuously run in parallel even if the virtual DCS safety operator 121 is processing an incident response. The virtual DCS safety operator 121 is multi-threaded and can handle multiple potential safety issues in parallel. The Kubernetes operator is typically implemented in the programming language Go (which is also used by Kubernetes), but can be implemented in any programming language. In its three phases, the virtual DCS safety operator 121 can act as follows according to several examples of the present disclosure:
[0093] Incident monitoring: The virtual DCS safety operator 121 permanently monitors data D1 Figure 4 , step 1) from the Kubernetes API and data D2 Figure 4 , step 1) from the registered OPC UA server. For example, the virtual DCS safety operator 121 can permanently monitor Kubernetes events and OPC UA events via a Kubernetes client 401 and an OPC UA client 402 comprised in the virtual DCS safety operator 121. The data D1, D2 can include specific device states, a list of processed alarms, typical audit trail data, pod life cycle, configuration changes, etc. For larger facilities, this can be supported by a distributed event streaming platform such as Apache Kafka.
[0094] Incident detection: The incident detector 403 is equipped with Kubernetes events and OPC UA events, can filter data, and applies incident detection rules 404 to the data or filtered data. Such new rules can be integrated immediately after a user 110 specifies a new rule in the virtual DCS security customization resource 122 (steps 3 and 4). The incident detector 403 can also be extended to identify patterns in the data and possibly suggest new rules by itself. Because the incident detector 403 has access to process data D2 and IT data D1, the incident detector 403 can correlate different event streams like Kubernetes events and OPC UA events and thus can potentially reveal additional security breaches that would otherwise go unnoticed. For example, unusual sensor readings found through OPC UA like a motor irregularly starting and stopping can be correlated with newly started pods in the system or unusual configuration changes in Kubernetes. This would not easily be possible if process data D2 and IT data D1 were analyzed separately. In case of an actual incident, the virtual DCS security operator 121 can inform the user 110 via the user interface 405 (step 5) and / or directly act by passing incident information to the incident responder 406. Figure 4
[0095] Incident response: The incident responder 406 operates similarly to the incident detector 403 according to custom incident response rules 407 specified as virtual DCS security customization resources 122. In addition to applying commands from the user interface 405 (step 6), the incident responder 406 can in some cases act autonomously and directly apply pre-specified incident response rules 407 without user interaction (steps 7 and 8). This allows for a fast reaction to security breaches. The incident responder 406 passes commands or incident response commands to the IT infrastructure (step 9). The incident responder 406 can utilize the entire Kubernetes API (step 10) and the connected OPC UA server (step 11) to issue incident response commands. For example, the incident responder 406 can redeploy a possibly compromised Kubernetes pod to a secure quarantine zone in the cluster. In more severe cases, the incident responder 406 can partially shut down all non-safety-critical pods in the system like supervisory DCS pods to contain a security incident. This is only possible because the incident responder 406 has incident response rules 407 that clearly identify non-safety-critical pods and cannot be done by a general-purpose incident response system. Figure 4 Figure 4 Figure 4 Figure 4 Figure 4
[0096] In more advanced variants, incident responder 406 can even attempt to "simulate" certain incidents before executing a response. Simulation can include a copy of a DCS pod or even a virtual pod, and can help assess the consequences of a partial system shutdown.
[0097] Incident responder 406 can also interactively attempt to formulate an appropriate incident response using a session user interface, iteratively feeding current incident detection information and cluster state into the Large Language Model (LLM). Hints to the LLM can request advice on how to handle the situation and even predict the consequences of a particular incident response.
[0098] Successful incident responses, including actions and commands, can be archived and turned into new incident response rules so that they can be quickly executed again in the future if a similar situation occurs.
[0099] According to several examples of this disclosure, a data processing apparatus is provided for incident detection in a cloud-local DCS during industrial process automation. The data processing apparatus can be configured to perform... Figure 2 Methods and / or references Figure 4 The method outlined (steps 1 to 11). The data processing device can be represented and / or used as a reference as described above. Figure 1 The virtual DCS security operator 121 described in 3 and 4.
[0100] More specifically, based on various examples, it is configured to execute Figure 2 Methods and / or execution Figure 4 The data processing device of the method may include a processing circuit system, processing function, processing apparatus, processing unit, or processor, which enables the data processing device to participate in incident detection in a cloud-local DCS in industrial process automation. The processor may include one or more processing portions or functions, wherein the processing portions or functions may be provided as one or more physical or virtual entities. The data processing device may include one or more communication interfaces. The data processing device may also include memory or storage units for storing data, programs, and / or instructions to be executed by the processing unit. The memory may be internal to the data processing device or external to the data processing device, such as at a cloud server. The processor may include one or more portions that enable the data processing device to perform, for example... Figure 2 The method. According to several examples of this disclosure, the monitoring section can be configured to perform actions based on... Figure 2 For monitoring like the S210, the joint analysis component can be configured to perform based on... Figure 2 The joint analysis of S220, for example, allows the detection component to be configured to perform based on...Figure 2 such detection of S230 and the response part can be configured to perform a response according to Figure 2 such response of S240.
[0101] For example, a part of the data processing device can also be understood to denote a means for performing a certain function.
[0102] According to several examples of the present disclosure, a data processing system for security incident detection in cloud-on-premises DCS in industrial process automation is provided. The data processing system can comprise a data processing device as described above, configured to perform the method of Figure 2 and / or to perform the method of Figure 4 Additionally or alternatively, the data processing system can be configured to perform the method of Figure 2 and / or to perform the method of Figure 4 The data processing system can be the contextual system 100 as described above with reference to Figure 1 .
[0103] According to several examples of the present disclosure, an industrial plant comprising a data processing device as described above and / or a data processing system as described above is provided.
[0104] According to several examples of the present disclosure, a computer-readable medium comprising instructions which, when executed by a computing system, cause the computing system to perform the method of Figure 2 and / or to perform the method of Figure 4 The computer-readable medium can be transitory or non-transitory, volatile or non-volatile.
[0105] According to several examples of the present disclosure, a computer program product comprising instructions which, when executed by a computing system, enable or cause the computing system to perform the method of Figure 2 and / or to perform the method of Figure 4 The computer program product can comprise a computer-readable medium comprising the instructions of the computer program product.
[0106] According to several examples of the present disclosure, a use of a data processing device as described above, and / or of a data processing system as described above, and / or of an industrial plant as described above is provided.
[0107] The method of Figure 2 and / or Figure 4 may be computer-implemented.
[0108] The method of Figure 2 and Figure 4Optional features of the methods described can form part of any data processing apparatus, data processing system, industrial plant, computer readable medium, computer program product, and use, mutatis mutandis.
[0109] Any of the units, modules, circuitry or methods described herein can be implemented using hardware, software, and / or firmware configured to perform the operations described herein. The hardware can include one or more processor cores, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc. The software can be implemented as software packages, code, instructions, instruction sets, and / or data recorded on at least one computer readable storage medium. The firmware can be implemented as code, instructions, or instruction sets and / or hard-coded data in a memory device (e.g., non-volatile memory device).
[0110] If implemented in software, the functions can be stored or transmitted over as one or more instructions or code on a computer readable medium. Computer readable media include computer readable storage media. A computer readable storage media is any available storage media that can be accessed by a computer. By way of example, and not limitation, such computer readable storage media can comprise FLASH memory, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Disk and disc, as used herein, includes compact discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks and Blu-ray discs (BD), where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Further, a propagated signal is also included in computer readable media. Computer readable media also includes communication media including any medium that facilitates the transfer of computer program from one place to another. A modulated data signal sent through a wired medium or wirelessly can be included as a communication media. By way of example, and not limitation, communication media includes wired media such as twisted pair wires, coaxial cables, and fiber optic cables, and wireless media such as acoustic, RF, infrared, and microwave. Combinations of the above should also be included within the scope of computer readable media.
[0111] The applicant hereby disclaims any right to a patent on an embodiment of the application merely because such embodiment is readily discovered from the disclosure or otherwise known to others in the art. The applicant hereby affirms that the aspects of the application can include any individual feature or combination of features disclosed herein.
[0112] It must be noted that embodiments of the application are described with reference to different classes of subject matter. In particular, some examples are described with reference to methods, while other examples are described with reference to devices. However, a person skilled in the art will gather from the description herein (and from the claims) that any combination of features from different classes of subject matter can also be possible and are to be considered disclosed herein as if each and every combination is explicitly claimed. All features available under a class of subject matter can be combined, provided this results in a technically functional combination.
[0113] While the application has been illustrated and described in detail in the drawings and foregoing description, such illustration and description is to be considered exemplary and not restrictive in character. The application is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in
[0114] The mere fact that measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0115] Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. A method for detecting safety incidents in a cloud-local distributed control system (DCS) for industrial process automation, the method comprising: Monitor (S210) information technology (IT) related data and operational technology (OT) related data at the production process and the containerized DCS associated with the production process; The joint analysis (S220) includes first data (D1) and second data (D2), wherein the first data indicates first monitoring data from the IT-related data, and the second data indicates second monitoring data from the OT-related data, and the joint analysis is based on associating at least a portion of the first data (D1) with at least a portion of the second data (D2) and / or based on associating at least a portion of the second data (D2) with at least a portion of the first data (D1); Based on the joint analysis, a safety incident is detected (S230) taking into account the predetermined safety incident detection rules (404); as well as Based on the detection results, a response (S240) is made to the detected security incident to process the detected security incident in consideration of a predetermined security incident response rule (407).
2. The method of claim 1, further comprising the monitoring, the joint analysis, the detection, and the response performed by a virtual DCS security operator (121), wherein the virtual DCS security operator (121) is a software agent running in a container orchestration cluster associated with the DCS; and / or wherein the virtual DCS security operator (121) is an autonomously operating security operator.
3. The method according to claim 1 or 2, Wherein the first monitoring data represents monitored IT-related data including monitored system diagnostic data, and wherein the second monitoring data represents monitored OT-related data including monitored process data; and / or The association includes associating at least a portion of the first monitoring data with at least a portion of the second monitoring data and / or associating at least a portion of the second monitoring data with at least a portion of the first monitoring data.
4. The method according to any one of claims 1 to 3, The monitoring of the IT-related data and the monitoring of the OT-related data include monitoring the IT-related data and the OT-related data while taking into account predetermined security incident monitoring rules; and / or The monitoring of the IT-related data and the monitoring of the OT-related data include monitoring data from the Kubernetes Application Programming Interface (API) server and the Open Platform Communication Unified Architecture (OPC UA) server; and / or The method also includes accessing the Kubernetes API server and executing the response based on adjusting parameters that can be used on the Kubernetes API server.
5. The method according to any one of claims 1 to 4, wherein the joint analysis comprises detecting a security incident in one of the first data (D1) and the second data (D2); and analyzing the other of the first data (D1) and the second data (D2) in relation to an event associated with the detected security incident.
6. The method according to any one of claims 1 to 5, The production process and the containerized DCS correspond to a specific domain, and The predetermined security incident detection rule (404) and the predetermined security incident response rule (407) are specific to the specific domain.
7. The method according to any one of claims 1 to 6, wherein the method further comprises using a virtual DCS security customization resource (122), the virtual DCS security customization resource (122) comprising at least a portion of the predetermined security incident detection rule (404), at least a portion of the predetermined security incident response rule (407), and at least a portion of the predetermined security incident monitoring rule.
8. The method according to claim 7, further comprising modifying the virtual DCS security customization resource (122) for at least one of the predetermined security incident detection rule (404), the predetermined security incident response rule (407), and the predetermined security incident monitoring rule. The modifications are performed manually by the user (110), automatically by the inference system associated with the DCS without involving the user (110), or semi-automatically when the user (110) provides guidance to the inference system.
9. The method according to any one of claims 1 to 8, wherein the response comprises at least one of the following: The detected security incident will be notified to the user (110); The application receives commands from the user (110) regarding the detected security incident; Autonomously apply the predetermined security incident response rules (407) regarding the detected security incident; as well as Before executing the response, a response to the detected security incident is simulated, and the response is further executed based on the results of the simulation.
10. The method according to any one of claims 1 to 9, wherein the method further comprises Exchange third-party data (D3) with the Security Information and Event Management (SIEM) system (310); and Further, at least one of the monitoring, the joint analysis, the detection, and the response is performed based on the third data (D3). The third data (D3) indicates at least one of the following: The recorded events that occur at the production process and / or the containerized DCS, The response executed Additional, removed, and / or updated security incident monitoring rules, Additional, removed, and / or updated security incident detection rules, and Additional, removed, and / or updated security incident response rules.
11. A data processing apparatus (121) for detecting safety incidents in a cloud-local distributed control system (DCS) in industrial process automation, the data processing apparatus comprising a processor configured to perform the method according to any one of claims 1 to 10.
12. The data processing apparatus (121) according to claim 11, comprising: Kubernetes client (401), OPC UA client (402), incident detector (403), incident responder (406), and user interface (405), The Kubernetes client (401) is communicatively connected to the incident detector (403) and the incident responder (406). The OPC UA client (402) is communicatively connected to the incident detector (403) and the incident responder (406). The incident detector (403) is communicatively connected to the user interface (405) and the incident responder (406), and The incident responder (406) is communicatively connected to the user interface (405) and the incident detector (403). The Kubernetes client (401) and the OPC UA client (402) are configured to perform the monitoring as described in claim 1. The accident detector (403) is configured to perform the joint analysis and detection as described in claim 1, and The incident responder (406) is configured to perform the response as described in claim 1.
13. A data processing system (100) for detecting safety incidents in a cloud-local distributed control system (DCS) for industrial process automation, the data processing system (100) comprising the data processing device (121) according to claim 11 and / or the data processing system (100) comprising means for performing the method according to any one of claims 1 to 10.
14. A computer-readable medium comprising instructions that, when executed by a computing system, cause the computing system to perform the method according to any one of claims 1 to 10.
15. A computer program product comprising a virtual DCS security operator (121) according to any one of claims 1 to 10, the computer program product comprising instructions that, when executed by a computing system, enable the computing system to perform and / or cause the computing system to perform the method according to any one of claims 1 to 10.