Methods for detecting system problems and methods for distributing micro mist
By deploying software agents and machine learning algorithms in the fog computing network to dynamically monitor and optimize micro-fog distribution, the problem of application distribution automation in the fog network is solved, and the system management efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202110719022.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-30
- Filing Date
- 2021-06-28
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-06-28
AI Technical Summary
Application allocation in fog computing networks needs to be improved, especially when equipment fails or the network changes. Existing technologies find it difficult to automatically and efficiently manage application allocation, resulting in network inefficiency.
By deploying software agents on system devices to monitor system configuration and functions, dynamically increasing or decreasing monitoring density, automatically detecting problems and collecting data, and combining machine learning algorithms to update rule groups, the distribution and deployment of micro-mist can be optimized.
It achieves more efficient and dynamic application allocation and management in fog networks, reduces system failures and downtime, and improves system observability and resource utilization.
Smart Images

Figure CN113867932B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to distributed control systems, fog computing networks, and particularly to managing fog applications in a fog network implemented on an automation system. Background Art
[0002] The advent of the Internet of Things (IoT) is extending the availability of networked computing and resources to a wide range of devices and systems previously excluded from data networking environments. Devices that once operated independently and were manually programmed can now work together and interact with one another. Complex systems comprise multiple devices working together as automated systems that react to and interact with their environment.
[0003] The goal is to achieve higher levels of automation by enabling machines of varying complexity and purpose to communicate without relying on human intervention and / or through manually programmed interaction of the machines via interfaces. Most devices, sensors, and actuators ("things") that are network-enabled in this way will typically be included in larger systems that provide new forms of automation. Industrial automation systems are becoming "smarter," and fog computing can help improve engineering efficiency.
[0004] Fog computing helps enable these larger systems by moving the computing, networking, and storage capabilities of a centralized cloud closer to machines and devices. Given the projected scale of such systems, demand for fog node resources is expected to be high.
[0005] Previously available cloud solutions (e.g., computing and storage) have many shortcomings and limitations that prevent them from meeting the performance requirements of IoT applications. For example, previously available cloud solutions lack the performance to meet low latency thresholds, support highly mobile endpoint devices, and provide real-time data analysis and decision-making.
[0006] Fog computing networks (or fog networks or fog environments) are being developed as a solution to meet the performance demands of IoT applications. Fog networks provide computing and storage resources closer to the edge of the network, rather than the remote and centralized clusters of previously available cloud solutions. The fog network's endpoint client devices and near-user endpoint devices are configured to collaboratively serve client applications at the edge of the network, close to the things seeking resources.
[0007] The Industrial Fog Environment enables easy deployment of fog applications on spare resources (so-called fog nodes) in the network and computing devices of industrial automation systems. To ensure that application components (so-called foglets) have sufficient available resources to complete their functions, resources are reserved for them during hardware sizing based on a declared estimated resource usage model. However, failures in devices or software components can lead to inefficiencies in the fog network.
[0008] Therefore, a concept for monitoring the execution of fog applications across fog networks is needed.
[0009] The distribution of applications within fog networks requires computation based on a model. Typically, the application model is manually adapted if the fog network changes or if an application is used for the first time on the network. Fog networks enable the running of distributed applications across devices in the underlying automation system. A key characteristic of fog computing is that application deployment, updates, and removal require minimal manual effort.
[0010] Therefore, there is a need to improve application distribution and automation in fog networks. Summary of the Invention
[0011] A method for detecting system problems in a distributed control system comprising a plurality of computing devices is provided. The method comprises: deploying one or more software agents on one or more devices of the system; monitoring system configuration and / or system functionality via the one or more software agents; detecting problems in the monitored system configuration and / or system functionality; adding one or more new software agents and deploying the one or more new software agents on one or more devices of the system associated with the problem; and collecting data associated with the problem via the added software agents.
[0012] A method for distributing mistlets in a fog network is proposed, wherein the fog network is implemented on a distributed control system comprising a plurality of devices. The method comprises: providing a distributed control system comprising a plurality of devices, wherein one or more devices provide computing capacity; providing a fog network implemented on the system having a plurality of fog nodes; providing a first group of mistlets comprising at least one mistlet; distributing the first group of mistlets to one or more fog nodes, wherein the distribution is based on a predetermined rule set for mistlet distribution; monitoring key performance indicators of the execution of the first group of mistlets; automatically creating or updating a dynamic rule set for mistlet distribution based on the distribution of the first group of mistlets and the monitored key performance indicators; providing a second group of mistlets comprising at least one mistlet and distributing the second group of mistlets to one or more fog nodes, wherein the distribution is based on the predetermined rule set for mistlet distribution and the dynamic rule set for mistlet distribution, or moving the execution of at least one mistlet in the first group of mistlets from one fog node to another fog node based on the predetermined rule set for mistlet distribution and the dynamic rule set for mistlet distribution.
[0013] Those skilled in the art will recognize additional features and advantages upon reading the following detailed description and upon viewing the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 An example of an automated system with multiple devices is shown. A fog network is hosted by the device;
[0015] Figure 2 An example of a method according to the present disclosure is shown; and
[0016] Figure 3 An example of a method according to the present disclosure is shown. DETAILED DESCRIPTION
[0017] In the following detailed description, reference is made to the accompanying drawings which form a part hereof, and in which is shown by way of illustration specific embodiments of the invention.
[0018] As used herein, the terms “having,” “containing,” “including,” “comprising,” and the like are open-ended terms that indicate the presence of stated elements or features, but do not preclude additional elements or features.
[0019] It should be understood that other embodiments may be used and that structural or logical changes may be made without departing from the scope of the present invention. Therefore, the following detailed description should not be considered restrictive, and the scope of the present invention is defined by the appended claims. The embodiments described herein use specific language, which should not be interpreted as limiting the scope of the appended claims. Unless expressly stated otherwise, each embodiment and each aspect so defined may be combined with any other embodiment or any other aspect.
[0020] Figure 1 An exemplary embodiment of a distributed control system (DCS) or distributed automation system or simply a distributed system implemented on multiple devices is shown. Figure 1 The DCS in [1] is a fog network. Fog computing is also known as edge computing or atomization. Fog computing facilitates the operation of computing, storage, and network services between terminal devices and cloud computing data centers. Some prior art documents describe the differences between fog computing and edge computing. However, this application does not distinguish between the concepts of fog computing and edge computing, and at least in the concept of this invention, considers the two to be the same. Therefore, any reference to fog computing can be a reference to edge computing, any reference to fog applications can be a reference to edge applications, and so on.
[0021] Physical systems can be automated systems. Any system that interacts with its environment, for example, via sensors or actuators, is considered an automated system. An automated system can include a large number of heterogeneous devices. These devices are sometimes referred to as "things," and the concept of things connecting and communicating with one another is also known as the "Internet of Things."
[0022] Figure 1 The diagram shows multiple devices that can form an automation system. It also shows a fog network, an optional remote service center, and a cloud. The fog network can be connected to the cloud. The cloud can provide additional computing resources, such as memory or CPU capacity. Data generated by the devices can be collected and transmitted to the cloud for analysis. While the cloud offers scalability benefits, pushing more and more computing from the automation system into the cloud has limitations and is often technically or economically unfeasible. Transmitting large amounts of data creates issues with latency thresholds, available bandwidth, and delayed real-time responses in the automation system.
[0023] A device may have computing resource capacity, such as CPU capacity, memory capacity, and / or bandwidth capacity. The resource capacity of some devices is Figure 1 exemplarily shown in FIG, to show a pie chart corresponding to the device, where the full resource capacity is a complete circle, the idle resource capacity is the unfilled portion of the circle, and the resource capacity in use is the filled portion of the circle.
[0024] Some devices are considered "smart" devices, while others are considered "dumb." Smart devices can host fog nodes and / or provide resource capacity. Some examples of smart devices are Industry 4.0 equipment, automated machines, robotic systems, user interfaces, network devices, routers, switches, gateway devices, servers, and similar devices. Some devices may not host fog nodes or may have very simple tasks, such as "read-only" devices or "zero resource capacity" devices. However, these "dumb" devices can still interact with the fog network, even though they do not provide resource capacity for tasks beyond their primary function, such as simple sensors, simple actuators, or other similar devices.
[0025] Figure 1 A simple example of an automation system or distributed control system (DCS) with a small number of devices is shown for illustrative purposes. However, other automation systems may include a greater number of devices. An automation system may typically be a factory or industrial site that includes any number of devices. The devices of an automation system may be heterogeneous and may provide different resource capacities.
[0026] A fog network consists of multiple fog nodes. Figure 1 The fog network in [1] is a network of fog nodes, which reside on many devices anywhere between the device farm and the cloud. Fog nodes provide the execution environment for the fog runtime and foglets. The fog runtime is the management application configured to authenticate foglets.
[0027] Furthermore, the fog network includes software distributed and executed on various components to manage the fog network and implement the functionality described below. The fog network can deploy and run potentially distributed fog applications. Based on the application model, it can determine which application parts (so-called foglets) should be deployed on which fog nodes. Therefore, it allocates application parts (foglets) to adhere to given constraints and optimize one or more objectives based on the application's requirements, as explained further below. Furthermore, the fog network may be able to incorporate fog nodes into the cloud, but should not be cloud-dependent.
[0028] Fog nodes are implemented on devices. A device can host one or more fog nodes. If fog applications are running on a device shared with non-fog applications, the fog applications should not interfere with these applications (the so-called primary functions). Therefore, fog nodes can use the predetermined maximum resource capacity of the hosting device. However, in some examples, the maximum resource capacity can also be variable. Fog nodes can be hosted in parallel with the primary functions of the device. In some embodiments, fog nodes can be hosted on virtual machines (VMs) on the device.
[0029] Figure 1 The "Fog Orchestration" shown is a conceptual illustration of the basic software components that can be included in a fog network and can run on one or more fog nodes. The software components can include:
[0030] A fog monitor that retrieves information about fog nodes and deployed fog applications. The resulting model of fog network and application allocation can be provided to one or both of the fog controller and fog manager;
[0031] A fog controller, including a fog distributor, computes a mapping of application parts (mistlets) to fog nodes such that data transmission is minimized. It then deploys the computed mapping and mistlets;
[0032] The fog manager provides a user interface for selecting and configuring fog applications, allows users to deploy / update / delete applications, and displays information about the fog network and deployed applications. The fog manager triggers corresponding functions of the fog controller based on the user's deployment / update / deletion requests.
[0033] According to one aspect of the present disclosure, a micromist is a deployment, execution, and management unit in a fog network. A fog application typically consists of a set of micromistlets, which together form value-added functionality. In other words, a micromist is an application component. A micromist refers to a specific function (e.g., a module) of an application and how it is deployed and executed, e.g., a micromist configuration. Such application modules are the building blocks of fog applications.
[0034] The allocation of mistlets can use an allocation algorithm that calculates the allocation of mistlets to fog nodes based on a specific application model. For example, the algorithm can implement a heuristic or exact linear program solution with a specific optimization goal as the target, such as minimizing the data transmission across the network link. Based on the application model and the fog network on which the application should be deployed, the allocation algorithm (allocator) calculates the mapping of application parts (mistlets) to fog nodes. Thus, the allocation algorithm can have multiple optimization goals, for example, it should minimize the required network bandwidth, minimize latency, meet network bandwidth constraints and constraints on data flow latency, and meet specific requirements indicated in the application model.
[0035] According to Figure 2 One aspect of the present disclosure provides a method for detecting system problems in a distributed control system comprising multiple computing devices. The method includes: deploying one or more software agents on one or more devices in the system; monitoring system configuration and / or system functionality via the one or more software agents; detecting problems in the monitored system configuration and / or system functionality; adding one or more new software agents and deploying the one or more new software agents on one or more devices in the system associated with the problem; and collecting data associated with the problem via the added software agents.
[0036] In a distributed control system (DCS), devices can join or leave the running system, for example, when the DCS is equipped with new smart sensors. Consequently, it becomes increasingly difficult to monitor all areas of the system from a few pre-selected fixed devices (e.g., gateways and firewalls). This means that it becomes increasingly difficult to detect system problems (e.g., security vulnerabilities and / or simple misconfigurations), which can lead to system failures and / or costly downtime. The proposed method addresses this problem by providing a software agent on one or more devices in the system that monitors system configuration and / or system functionality.
[0037] The present disclosure mitigates such problems by adaptively distributing software agents strategically across the system, which always provides the required level of system visibility. If a suspected problem arises (a problem in the monitored system configuration and / or system functionality), the number of monitoring software agents in the affected area is dynamically increased or decreased based on the current observability requirements. For example, when reconfiguring a portion of the system, the agent density can be temporarily increased to detect any misconfigurations in that portion of the system. Thus, the present invention enables a more dynamic system with less overall monitoring effort and therefore at a lower cost.
[0038] According to one aspect, monitoring of system configuration and / or system functionality includes monitoring of network traffic, application data, system performance, or a combination thereof. A set of active software agents constitutes a collaborative and distributed system that processes network traffic, application / network data, or system performance data in real time.
[0039] According to one aspect, the monitoring of the system is an on-the-fly observation of the running system, particularly in terms of the depth or breadth of the running system. The monitoring can be performed periodically in near real time, for example, every second, every two seconds, or every five seconds.
[0040] According to one aspect, detecting problems in the monitored system configuration and / or system functionality includes: - comparing the monitored system configuration and / or system functionality with a known or expected system configuration and / or system functionality. The changed system configuration may be caused by a change in system topology, such as devices joining and leaving the system. Alarms or events may be recorded by comparing the monitored system configuration and / or system functionality (the current state of the system) with a known / expected / anticipated system model.
[0041] The data associated with the problem collected via the added software agent can be the same type of data that is monitored during monitoring of system configuration and / or system functionality, but in greater detail. For example, a software agent can monitor network traffic at a gateway and detect a problem. A new software agent can then be sent to devices connected to the gateway to determine which device behind the gateway has the problem.
[0042] The known or expected system configuration and / or system functionality may also be a defined system configuration and / or system functionality based on a (previously) monitored system configuration and / or system functionality. The method may comprise: defining a normal system configuration and / or normal system functionality based on the monitored system configuration and / or system functionality, wherein detecting a problem in the monitored system configuration and / or system functionality comprises detecting a problem in the monitored system configuration and / or system functionality by comparing the monitored system configuration and / or system functionality with the normal system configuration and / or normal system functionality.
[0043] According to one aspect, normal system configuration and / or normal system functionality can be continuously updated based on monitoring. For example, minor changes in system configuration and / or system functionality can be considered "normal," which will result in an update of the normal system configuration and / or normal system functionality. A sudden change in system configuration or a sustained decrease in functionality can be considered a "detected problem."
[0044] Specifically, updating normal system configuration and / or normal system functionality can be based on machine learning algorithms. The software agent can "learn" the normal system configuration and / or normal system functionality of a given system based on monitoring data during or before operation. Problem detection can also be based on machine learning algorithms.
[0045] In some examples, the monitored issues in system configuration and / or system functionality are associated with adding or removing devices from the system.
[0046] A software agent is a software component configured to monitor system configuration and / or system functionality. Specifically, a software agent can be a fog micro-monitoring application in a fog network. The number of active software agents running across devices or fog nodes in a running system can depend on the total number of devices in the system and / or the desired monitoring density.
[0047] In some examples, software agents are provided in a software agent repository and the method may include: providing a software agent repository comprising a plurality of software agents, wherein deploying the one or more software agents on one or more devices of a system comprises selecting one or more software agents from the plurality of software agents in the software agent repository and deploying the selected one or more software agents on the one or more devices of the system. Each software agent may be of a certain type. The types of software agents may be configured to connect to different data sources and execute monitoring queries as complex event processing.
[0048] According to one aspect, the method may further include: reporting the collected data; and deleting the added software agent. After receiving the data needed to identify the exact problem, an engineer can use the data to manually resolve the issue. After collecting the data, the software agent can be removed or decommissioned.
[0049] A software agent can be described as a standardized unit of software that can be stored in a repository, referred to as a "static software agent," or downloaded to a device belonging to a system and executed, referred to as a "runtime software agent."
[0050] The set of running software agents is scalable and adaptable to the state / condition of the system. In one operating mode ("supervision"), the software agents can perform high-level monitoring using a minimal number of agents and / or reduced agent functionality to reduce overhead. In another operating mode ("deep inspection"), new agents can be added temporarily (or existing agents can be duplicated) to investigate possible issues in the system that require increased monitoring to better understand the problem. Once executed, the software agent consumes a predetermined amount of resources on the physical device assigned to it.
[0051] According to one aspect, deploying one or more software agents on one or more devices of a system includes: - deploying at least a coordinator software agent configured to create and deploy software agents on one or more devices of the system; - creating and deploying one or more software agents on one or more devices of the system by the coordinator software agent, wherein the created and deployed software agents are configured to report to the coordinator software agent the resource requirements of additional software agents and / or whether the software agents should be terminated.
[0052] The coordinator software agent can be configured to terminate software agents on one or more devices in the system, wherein the method further comprises: terminating, by the coordinator software agent, one of the one or more software agents. Specifically, the coordinator software agent manages the lifecycle of software agents and their dispatch or replication to different areas of the system to increase observability in those areas. The method includes a mechanism that allows existing "running software agents" to communicate to the coordinator software agent resource requirements for new software agents or whether certain "running software agents" should be decommissioned.
[0053] The orchestrator software itself may be replicated for redundancy and high availability. Thus, deploying at least an orchestrator software agent configured to create and deploy software agents on one or more devices of the system may comprise: - deploying a plurality of orchestrator software agents, each orchestrator software agent configured to create and deploy software agents on one or more devices of the system.
[0054] In some examples, the software agent can be a fog agent, which is a type of micro-mist. A fog network including a plurality of fog nodes can be implemented on a system, wherein deploying the software agent on one or more devices of the system includes deploying the software agent (fog agent) on one or more fog nodes among the fog nodes implemented on the one or more devices.
[0055] A system is also proposed, wherein the system comprises a plurality of devices and wherein the system is configured to perform the method as described herein.
[0056] The following examples illustrate some embodiments of the present disclosure:
[0057] Example 1
[0058] In this example, the system includes software agents whose role is to detect anomalies in network traffic while keeping overhead low by deploying specialized agents only when needed. By default, the software agents observe and follow simple communication patterns, so that the computational overhead introduced by these agents remains low. Once the software agents observe that certain communications do not conform to previously learned patterns, for example, a sudden increase in the use of a specific operational technology (OT) communication protocol, a deeper inspection is required. To this end, a set of stationary software agents dedicated to inspecting specific OT communication protocols are now awakened and deployed on devices (or fog nodes) that provide better observability of the problem. The goal is to assess whether the observed behavior is indeed an intrusion. Once the inspection is completed, these software agents are removed from their physical devices and become "stationary fog agents" again.
[0059] Example 2
[0060] In this example, the system includes software agents whose role is to detect anomalies in network traffic, with the goal of early detection and localization of intrusions. This is done by maintaining low overhead by regulating system observability through the number and location of deployed agents. To reduce overhead, agents are deployed at network aggregation points within the system, such as edge routers, to observe all traffic passing between different parts of the system. By observing traffic, the agents learn normal traffic patterns. Once an expected change in traffic patterns from a particular part of the system occurs—for example, a sudden increase in traffic indicating a denial of service attack or some system nodes suddenly using a new communication protocol—a deeper system inspection is required to learn more about the anomaly and locate potential intrusions. A new set of software agents is then created or awakened from a quiescent state and deployed deep into that part of the system to identify the device causing the anomaly and investigate whether an intrusion has occurred. Once the investigation is complete, these software agents are removed from the physical device and become "quiescent software agents" or are removed.
[0061] Example 3
[0062] In Example 3, the system is a fog system in supervised mode, and software agents in the form of Network Anomaly Detection Sensor (NADS) agents are placed on fog nodes so that at least one of the two endpoints of each Ethernet connection is covered. The agent monitors all network traffic at its location and regularly evaluates (e.g., every second) whether the current system functionality is normal. Whether the situation is normal is evaluated based on a normal model that is either designed based on the system configuration or learned by machine learning methods during operation or before (e.g., during debugging). In supervised mode, the normal model is based on simple key performance indicators (KPIs) such as the number of packets per second and the number of different source-destination pairs (in the Ethernet and IP headers).
[0063] If the NADS agent detects an anomaly, it notifies the coordinator software agent. The coordinator software agent then initiates a deep inspection mode, i.e., it calculates a set of additional software agents and where they should be placed. In particular, NADS+ agents and DADS (Device Anomaly Detection Sensor) agents are placed at the device or fog node where the anomaly was detected, as well as at each adjacent node. NADS+ agents perform deeper network anomaly detection than NADS agents, i.e., they check more network service KPIs and can also perform protocol-specific deep packet inspection. DADS checks the logs available on the device, such as security-related system logs, to calculate cs. NADS+ and DADS agents report to the coordinator software agent. The coordinator software agent uses the more detailed information to perform root cause analysis and issue measures to mitigate the situation (e.g., disconnecting the infected device from the network) and / or provide a detailed report that can ultimately be used by human supervisors (via the device HMI, or via email or notifications in some kind of dashboard) to take appropriate action.
[0064] In a variation of Example 3, instead of adding separate NADS+ agents in addition to the NADS agents, the coordinator software agent may be able to reconfigure the NADS agents to upgrade them to NADS+ agents.
[0065] According to Figure 3Another aspect of the present disclosure is a method for distributing mistlets in a fog network, wherein the fog network is implemented on a distributed control system comprising a plurality of devices. The method comprises: providing a distributed control system comprising a plurality of devices, wherein one or more devices provide computing capacity; providing a fog network implemented on the system having a plurality of fog nodes; providing a first group of mistlets comprising at least one mistlet; distributing the first group of mistlets to one or more fog nodes, wherein the distribution is based on a predetermined rule set for mistlet distribution; monitoring key performance indicators of the execution of the first group of mistlets; automatically creating or updating a dynamic rule set for mistlet distribution based on the distribution of the first group of mistlets and the monitored key performance indicators; providing a second group of mistlets comprising at least one mistlet and distributing the second group of mistlets to one or more fog nodes, wherein the distribution is based on the predetermined rule set for mistlet distribution and the dynamic rule set for mistlet distribution, or moving the execution of at least one mistlet in the first group of mistlets from one fog node to another fog node based on the predetermined rule set for mistlet distribution and the dynamic rule set for mistlet distribution.
[0066] In the prior art, distributed control system (DCS) resources (e.g., CPU capacity, memory, storage) are statically allocated to specific functions, regardless of the resource requirements of the actual services (micro-mist). As a result, the performance of some functions may be affected by limited resources, while other functions may not use their allocated resources at all. The method disclosed herein addresses this problem by introducing two sets of rules on which micro-mist allocation is based.
[0067] A predefined set of rules for mist allocation can be provided during system engineering, possibly in conjunction with statically defined default allocations. Machine-readable rule definitions can be embodied, for example, as text files via a domain-specific scripting language. Simple examples of rules focus on the minimum resources required to run a particular service (e.g., a control execution service might require a certain amount of memory), or prerequisites from other services (e.g., partial execution).
[0068] In some examples, the method further comprises: - monitoring the current resource state of one or more fog nodes; wherein the allocation of the first group of mistlets is further based on the current resource state of the one or more fog nodes, and wherein the execution of allocating the second group of mistlets or moving the first group of mistlets is further based on the current resource state of the one or more fog nodes. The rules in the predetermined rule set for mistlet allocation and the dynamic rule set for mistlet allocation may include resource availability of the current resource state of the system. Thus, in this case, the allocator is able to obtain the resource state of one or each node (e.g., the node's CPU, memory, disk space). For example, such information may be provided by a fog monitor.
[0069] Specifically, the resource status includes one of the following: idle CPU capacity, idle memory, or a combination thereof.
[0070] Key performance indicators of the system can be performance data of network services or application data. For example, a key performance indicator (KPI) can be the number of packets sent per second or the number of unique source-destination pairs (in Ethernet and IP headers) in any type of network. A key performance indicator can also be the time required to execute an application, a group of mistlets, or a single mistlet.
[0071] This dynamic rule set is defined and can be updated, which enables the allocation algorithm to recognize the changing conditions of the system and "learn" the best way to distribute the mists in the system. This dynamic rule set optimizes the load by moving the execution of certain mists from one node to another (always complying with given predetermined rules).
[0072] According to one aspect, the dynamic rule set is automatically created or updated by an artificial intelligence algorithm based on key performance indicators monitored during the execution of the first set of micro mists. Specifically, the artificial intelligence algorithm can be a machine learning algorithm. The algorithm can improve the dynamic rule set over time during system operation.
[0073] The second group of mist particles is allocated based on the predetermined rule set for mist particle allocation and the dynamic rule set for mist particle allocation. The second group of mist particles may have one or more mist particles of the same type as the first group of mist particles.
[0074] According to one aspect, the method can be used in a continuously operating system. The first and second groups of mists can be continuously provided and distributed or repositioned based on the predetermined rule set for mist distribution and the dynamic rule set for mist distribution, wherein key performance indicators of mist execution are continuously monitored, and wherein the dynamic rule set is continuously and automatically updated based on the distribution of the first and second groups of mists and the monitored key performance indicators.
[0075] A system is also proposed, wherein the system comprises a plurality of devices and wherein the system is configured to perform the method as described herein.
[0076] Example
[0077] Services (mistlets) can be temporary in nature, i.e., they should be allocated and deployed to the system at a given point in time and removed again once their task is completed. For such services, it is crucial to have a mechanism that can dynamically reconfigure the system to provide the appropriate resources at the appropriate locations in the DCS, for example, through a combination of pre-defined and dynamic rule-based reallocation of other services exposed.
[0078] An example of this type of temporary service is a distributed engineering service. These engineering services may be needed by humans (system engineers) or other services. For example, a compute service may need to know when to start the compute frequency (which can be re-designed by a central engineering server). Thanks to the disclosed method, a system can generate an engineering service containing this information when necessary, deploy the required information on the nodes running our runtime service, and purge the information when it is no longer needed.
Claims
1. A method for detecting a system problem in a distributed control system comprising a plurality of computing devices, the method comprising: - deploying one or more software agents on one or more devices of the system; - monitoring system configuration and / or system functionality via said one or more software agents; - Detect problems in the monitored system configuration and / or system functionality; - adding one or more new software agents and deploying the one or more new software agents on one or more devices of the system associated with the problem; - collecting data associated with the problem via the added software agent, Wherein deploying one or more software agents on one or more devices of the system comprises: - deploying at least a coordinator software agent configured to create and deploy software agents on one or more devices of said system, - creating and deploying one or more software agents on one or more devices of the system by the coordinator software agent, wherein the created and deployed software agents are configured to report to the coordinator software agent the resource requirements of additional software agents and / or whether a software agent should be terminated. 2 . The method according to claim 1 , wherein the monitoring of the system configuration and / or the system function comprises monitoring network traffic, application data, system performance, or a combination thereof.
3. The method according to any one of the preceding claims, wherein detecting problems in the monitored system configuration and / or system functionality comprises: - Comparing the monitored system configuration and / or system functionality with known or expected system configuration and / or system functionality.
4. The method according to claim 1 or 2, further comprising: - Provide a software agent library including multiple software agents, Deploying one or more software agents on one or more devices of the system includes: selecting one or more software agents from the plurality of software agents in the software agent library, and deploying the selected one or more software agents on one or more devices of the system.
5. The method of claim 1 , wherein the coordinator software agent is further configured to terminate software agents on one or more devices of the system, wherein the method further comprises: - terminating, by the coordinator software agent, one of the one or more software agents.
6. The method of any one of claims 1 or 5, wherein deploying at least a coordinator software agent configured to create and deploy software agents on one or more devices of the system comprises: - deploying a plurality of orchestrator software agents, each orchestrator software agent being configured to create and deploy software agents on one or more devices of the system.
7. The method according to claim 1 or 2, wherein a fog network comprising a plurality of fog nodes is implemented on the system, wherein deploying a software agent on one or more devices of the system comprises: A software agent is deployed on one or more of the fog nodes implemented on the one or more devices.
8. The method of claim 1 or 2, wherein the problem in the monitored system configuration and / or system functionality is associated with adding a device to the system or removing a device from the system.
9. The method according to claim 1 or 2, further comprising: - defining a normal system configuration and / or normal system functionality based on the monitored system configuration and / or system functionality; Problems in the monitored system configuration and / or system functionality include: Problems in the monitored system configuration and / or system functionality are detected by comparing the monitored system configuration and / or system functionality with the normal system configuration and / or the normal system functionality.
Citation Information
Patent Citations
Method, Device and Computer Program for Monitoring an Industrial Control System
US20150301515A1