Distributed software-defined industrial system

By utilizing distributed software-defined industrial systems (SDIS) with dynamic data models and self-describing modules, the complexities of flexible updates and management of industrial systems in harsh environments are solved, enabling efficient and reliable system updates and resource optimization, and supporting efficient management of heterogeneous equipment and integration of new technologies.

CN111164952BActive Publication Date: 2025-10-21INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201880063380.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-12-29
Filing Date
2018-09-28
Publication Date
2025-10-21
Estimated Expiration
2038-09-28

AI Technical Summary

Technical Problem

Existing industrial systems struggle to adapt to harsh environments with varying temperatures, vibrations, and humidity, leading to increased management complexity, high costs, and a failure to fully leverage the potential of IoT and software-defined technologies.

Method used

The Distributed Software-Defined Industrial System (SDIS) enables flexible configuration and management of equipment through dynamic data models, orchestration technology, and self-describing modules. It supports efficient orchestration and resource optimization of heterogeneous equipment, manages alarms using intelligent machine learning, and improves system reliability through multi-layer field device redundancy buses.

Benefits of technology

It enables flexible system updates in harsh environments, reduces management complexity and costs, improves system reliability and computing resource adaptability, and supports efficient management of heterogeneous devices and rapid integration of new technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111164952B_ABST
    Figure CN111164952B_ABST
Patent Text Reader

Abstract

Various systems and methods for implementing software-defined industrial systems are described herein. For example, an orchestrated system with distributed nodes can run applications, including modules implemented on the distributed nodes. In response to a node failure, a module can be redeployed to a replacement node. In examples, self-describing control applications and software modules are provided in the context of an orchestratable distributed system. A self-describing control application can be executed by an orchestrator or similar control device and use a module manifest to generate a control system application. For example, an edge control node of an industrial system can include a system-on-chip that includes a microcontroller (MCU) for converting IO data. The system-on-chip includes a central processing unit (CPU) that is in an initial inactive state that can be changed to an active state in response to an activation signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority claim

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 62 / 587,227, filed on November 16, 2017, entitled “DISTRIBUTED SOFTWARE DEFINED INDUSTRIAL SYSTEMS,” and U.S. Provisional Patent Application Serial No. 62 / 612,092, filed on December 29, 2017, entitled “DISTRIBUTED SOFTWARE DEFINED INDUSTRIAL SYSTEMS,” which are incorporated herein by reference in their entireties. Technical Field

[0003] The embodiments described herein relate generally to data processing and communication within distributed and interconnected device networks, and more particularly to techniques for defining the operation of software-defined industrial systems (SDIS) provided from configurable Internet of Things devices and device networks. Background Art

[0004] Industrial systems are designed to capture real-world instrument (e.g., sensor) data and execute responses in real time while operating reliably and safely. The physical environment in which such industrial systems are used can be harsh and subject to wide variations in temperature, vibration, and humidity. Small changes to the system design can be difficult to implement because many statically configured I / O and subsystems lack the flexibility to be updated within the industrial system without completely shutting down the equipment. Over time, the incremental changes required to properly operate an industrial system can become overly complex and result in significant management complexity. Additionally, many industrial control systems incur expensive operating and capital expenditures, and the architecture of many control systems is not designed to take advantage of the latest advances in information technology.

[0005] The development of Internet of Things (IoT) technologies and software-defined technologies such as virtualization has led to technological advancements in many forms of telecommunications, enterprise, and cloud systems. Advances in real-time virtualization, high-availability, and secure software-defined systems and networking have provided improvements in such systems. However, IoT devices can be physically heterogeneous, and their software may also be heterogeneous (or may become increasingly so over time), complicating the management of such devices.

[0006] Despite technological advancements in industrial automation and systems, only limited approaches have been explored to leverage IoT devices and IoT frameworks. Furthermore, industry has been reluctant to adopt new technologies in industrial systems and automation due to their high cost and unproven reliability. This reluctance means that only incremental changes are typically attempted; even then, there are numerous examples of new technologies performing poorly or taking a long time to go live. As a result, large-scale deployment of IoT and software-defined technologies has yet to be successfully adapted to industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In the drawings (which are not necessarily drawn to scale), like numerals may describe similar components in different views. Like numerals with different letter suffixes may represent different instances of similar components. Some embodiments are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which:

[0008] Figure 1A Figure illustrates the configuration of a software defined infrastructure (SDIS) operating architecture according to a first example;

[0009] Figure 1B illustrates the configuration of an SDIS operating architecture according to a second example;

[0010] Figure 2A The diagram shows that according to an example Figure 1A Configuration of real-time advanced computing subsystems deployed within the SDIS operational architecture;

[0011] Figure 2B The diagram shows that according to an example Figure 1A Configuration of the edge control node subsystem deployed within the SDIS operating architecture;

[0012] Figure 3A The diagram shows that according to an example Figure 1B Configuration of real-time advanced computing subsystems deployed within the SDIS operational architecture;

[0013] Figure 3B and Figure 3C The diagram shows that according to an example Figure 1B Configuration of cloud computing and edge computing subsystems deployed within the SDIS operational architecture;

[0014] Figure 4 Figure illustrates the configuration of a control message bus used within the SDIS operating architecture according to an example;

[0015] Figure 5A FIGURE 1 illustrates a first network configuration for deploying an SDIS subsystem according to an example;

[0016] Figure 5BFIGURE 1 illustrates a second network configuration for deploying the SDIS subsystem according to an example;

[0017] Figure 6 illustrates protocols in an example scenario for dynamically updating a data model in an SDIS operational architecture according to an example;

[0018] Figure 7 illustratively a flow diagram for generating and utilizing a dynamically updated data model in an SDIS operational architecture according to an example;

[0019] Figure 8 illustrating a flow diagram of a method for incorporating a dynamically updated data model into use with an SDIS operational architecture, according to an example;

[0020] Figure 9 Figure illustrates a set of orchestration operations dynamically established in an SDIS operations architecture according to an example;

[0021] Figure 10 Figure illustrates an orchestration arrangement of a cascade control application based on distributed system building blocks according to an example;

[0022] Figure 11 FIGURE 1 illustrates an application distribution map of a control strategy for orchestrating a scenario according to an example;

[0023] Figure 12 Figure illustrates an orchestration scenario suitable for handling function block application timing dependencies according to an example;

[0024] Figure 13 illustrates orchestrating asset deployment for an application under the control of an orchestrator according to an example;

[0025] Figure 14 FIGURE 1 illustrates a flow diagram of an orchestrated sequence for distributed control application policies according to an example.

[0026] Figure 15 illustratively showing a flow diagram of a method for orchestrating distributed mission-critical workloads and applications using distributed resource pools according to an example;

[0027] Figure 16A FIGURE 1 illustrates a scenario of orchestration between an orchestration engine and associated modules according to an example;

[0028] Figure 16B FIGURE 1 illustrates a scenario of orchestration between an orchestration engine and associated modules including legacy modules according to an example;

[0029] Figure 17A illustrates a scenario with orchestration of orchestratable devices according to an example;

[0030] Figure 17Billustrates a scenario with an orchestration of legacy devices according to an example;

[0031] Figure 18 illustrates a coordinated scenario of workload orchestration in a single-level orchestration environment according to an example;

[0032] Figure 19 Figure illustrates a functional hierarchy of an arrangement according to an example;

[0033] Figure 20 illustrates deployment of a general layered orchestration solution according to an example;

[0034] Figure 21 Figure illustrates a hierarchical orchestration provided with use of slave nodes according to an example;

[0035] Figure 22 illustrates a workflow for a slave node used in a hierarchical orchestration scenario according to an example;

[0036] Figure 23 illustrates a configuration of a monitoring and feedback controller suitable for coordinating and implementing orchestrated self-monitoring functionality according to an example;

[0037] Figure 24 FIGURES a flow diagram illustrating an example method for orchestrating devices in a conventional setting according to an example;

[0038] Figure 25 The figure illustrates an industrial control application scenario according to an example;

[0039] Figure 26 illustrates an overview of a control application represented by a control application graph according to an example;

[0040] Figure 27 illustrates a self-describing software module definition for implementing a control application according to an example;

[0041] Figure 28 Figure illustrates an architecture for automatically evaluating alternative implementations of software modules according to an example;

[0042] Figure 29 illustratively a flow chart of a method for evaluating alternative implementations of a software module according to an example;

[0043] Figure 30A FIGURES illustrative of a flow chart of a method for implementing a self-describing programmable software module according to an example;

[0044] Figure 30B FIGURES a flow diagram illustrating a method for using self-describing programmable software modules in an SDIS system implementation according to an example;

[0045] Figure 31FIGURES illustrate a PLC-based industrial control system according to an example;

[0046] Figure 32 Figure illustrates a multi-layer field device bus according to an example;

[0047] Figure 33 Figure illustrates the IO converter functionality according to an example;

[0048] Figure 34 Figure illustrates IO converter redundancy according to an example;

[0049] Figures 35A-35B FIGURES illustratively show a flow chart of a method for implementing a multi-layer field device bus according to an example;

[0050] Figure 36 FIGURES AN example of a process with generated alerts according to an example;

[0051] Figure 37 illustrates a dynamic smart alert according to an example;

[0052] Figure 38 FIGURES a flow chart illustrating a method for dynamic alarm control according to an example;

[0053] Figure 39 The autonomous control-learning integration process is illustrated with an example diagram;

[0054] Figure 40 illustratively a flow chart of a method for managing autonomous creation of new algorithms for an industrial control system according to an example;

[0055] Figure 41 The figure shows a ring topology diagram of an industrial control system;

[0056] Figure 42 The figure shows the edge control topology;

[0057] Figure 43 The figure shows the edge control node topology;

[0058] Figure 44 The figure shows a ring topology based on edge control nodes;

[0059] Figure 45 The diagram illustrates the data flow through a ring topology based on edge control nodes;

[0060] Figure 46A FIGURES a flow chart illustrating a method for activating a processor of an edge control node according to an example;

[0061] Figure 46B FIGURES a flow chart of a method for activating a CPU according to an example;

[0062] Figure 47 The figure shows an example application connection diagram;

[0063] Figure 48 An example architecture diagram of an application with a standby node is shown;

[0064] Figure 49A illustratively showing a flow chart of a method for creating an automatic redundancy module of an application on redundant nodes based on a communication pattern of the application according to an example;

[0065] Figure 49B FIGURES a flow chart of a method for activating a CPU according to an example;

[0066] Figure 50 illustrates a domain topology for various Internet of Things (IoT) networks coupled to respective gateways via links according to an example;

[0067] Figure 51 illustrates a cloud computing network in communication with a mesh network of IoT devices operating as fog devices at the edge of the cloud computing network according to an example;

[0068] Figure 52 A block diagram illustrating a network according to an example illustrating communications between a large number of IoT devices; and

[0069] Figure 53 The diagram illustrates a block diagram of an example IoT processing system architecture on which any one or more of the techniques (eg, operations, processes, methods, and methodologies) discussed herein may be performed. DETAILED DESCRIPTION

[0070] In the following description, methods, configurations, and related apparatuses for configuring, operating, and adapting Software-Defined Industrial Services (SDIS) deployments are disclosed. Specifically, the following SDIS deployments include the functionality of industrial systems based on modern operating architectures, as well as derived architectures or solution instances of such deployments. For example, such architectures and instances may include virtualized control server systems that implement features of edge control devices and control message buses within control or monitoring systems. Such architectures and instances may further be integrated with various aspects of IoT networks, involving various forms of IoT devices and operations.

[0071] The processing techniques and configurations discussed herein include a variety of approaches for managing operations, data, and processing within various types of SDIS architectures. An overview of these approaches is provided in the following paragraphs; further references to specific implementation examples and use cases are discussed below.

[0072] In an example, a dynamic data model is established to provide a dynamic feature set for applications, devices, or sensors in an SDIS architecture. Such a dynamic data model can be data-driven in nature and can be contrasted with the statically defined data models typically established during development. For example, a dynamic data model can be represented by a device as a whole set of sensors, allowing the device to express itself using different output sensors based on changing factors (such as battery and computing availability). This dynamic data model can play an important role in making various systems and data in the Internet of Things available while being adaptable. The characteristics of the dynamic data model provide the device with the ability to be modified and extended at runtime, or even restored to a subset of its components. In addition, the dynamic data model can be embodied by dynamic metadata, complex representations of values ​​(including providing probabilistic estimates of labels rather than binary on / off states).

[0073] Also in the example, a configuration can be established in the SDIS architecture to support the overall orchestration and management of multiple dependent applications (e.g., function blocks) executed across a distributed resource pool. By including additional application-specific dependencies in the extended orchestrator logic rule set, orchestration can be enabled at the embedded control policy level in the distributed system configuration. Through dynamic discovery of network bandwidth, evaluation of resource capacity and current status, historical information and control application constraints, and similar information, a variety of multi-level optimization and prediction methods can be performed to complete advanced orchestration scenarios. Using such features, real-time events and predictions can also be used to hierarchical reactions into orchestration events to maintain the online status of a wider range of control policies. In addition, prediction and constraint management combined with real-time optimization of such orchestration can achieve a high level of resilience and functionality of the embedded infrastructure.

[0074] Also in the example, orchestration of functionality can be extended for existing forms of brownfield environments (where such "brownfield" devices refer to existing device configuration architectures). Orchestration in such traditional settings can be enabled by the use of underlays at both the application and device levels to support orchestration of unconscious application components and legacy devices; the use of a layered structure to support scale and legacy devices; and the adaptation of self-monitoring to manage the heterogeneity, resource utilization, scale, and built-in self-dependencies of various devices. The application of such orchestration technology in an SDIS architecture can be used to increase the scalability of the architecture to include many forms of devices, systems, and industries. Additionally, such orchestration technology allows the technology to be applied in situations where customers have already made significant investments in existing technology platforms.

[0075] Also in the example, the orchestration of functionality can serve as a key control point through which customers can leverage the differentiated capabilities of hardware deployments. This orchestration can be achieved through self-describing modules that provide a deployable mechanism for using self-describing control applications and software modules in the context of an orchestrated distributed system. Such self-describing modules allow for trade-offs between implementations, such as enabling customers to effectively use platform features when such features are available, while having alternatives when such features are not available. The following examples include implementations in an SDIS architecture suitable for automatically evaluating these trade-offs, thereby allowing for more efficient development of features for industrial use cases and deployments.

[0076] Also in the examples, the systems and methods described herein include a multi-layered field device redundancy bus that enables an "any-to-any" relationship between controllers and field devices. The decoupling of controllers and I / O enables simple failover and redundancy. Improved system reliability and survivability are achieved by enabling any controller to access any field device in the event of a controller failure. Reduced system costs can also be beneficial, such as by enabling the addition of new field devices based on small incremental investments rather than a heavy PLC burden.

[0077] Also in an example, the systems and methods described herein can use intelligent machine learning methods to manage alerts. The systems and methods described herein can: characterize data to detect anomalies that may trigger alerts; cluster alerts using data similarity or common causal relationships so that they are presented as a bundle to combat alert flooding and fatigue; or understand human responses to alerts so that these actions can be automatically performed in the future.

[0078] Also in the example, this paper proposes a rigorously sequenced policy framework and a series of methods for managing the autonomous creation of new closed-loop workloads in mission-critical environments through the following eight processes: quality and sensitivity assessment of new algorithms relative to processes; automated establishment of operational constraint boundaries; automated safety assessment of new algorithms relative to existing processes; automated value assessment with respect to the broader process; automated system assessment of deployment feasibility within the control environment; physical deployment and monitoring of newly applied control policies; integration into lifecycle management systems; and integration into end-of-life processing.

[0079] Also in the examples, the systems and methods described herein address the problem of over- or under-provisioning computing power at the edge of an industrial control system. Over-provisioning computing resources wastes money, electricity, and heat. Under-provisioning computing resources sacrifices reliability and the ability to execute control strategies. The proposed solution enables end users with performance requirement data to "right-size" the amount of computing to provision in a control environment. Additionally, the provisioned computing power is not static and can be adapted to meet the needs of the control system as demand changes. The techniques discussed herein allow a high-performance CPU to be activated from an initially dormant state in an edge control node using a centralized orchestration system that understands the CPU performance requirements of a control strategy.

[0080] Also in the example, additional module interconnection technologies are disclosed. In an orchestrated system, an application is typically defined as a set of modules interconnected by a topology. These modules are deployed on different logical nodes. Each logical node can correspond to a physical node, but the mapping does not have to be 1:1. As long as resource requirements are met, multiple logical nodes can potentially be mapped to one physical node, and multiple modules can be deployed on the same physical environment. In the example, the solution can create automatic backup nodes for modules based on the application's communication pattern. A peer-to-peer network created by a collection of nodes on the same layer can negotiate the status of backups. The community of nodes can also exchange backups among themselves without major impact on the rest of the application.

[0081] Other examples will be apparent from the following figures and textual disclosure.

[0082] Overview of Industrial Automation Systems

[0083] Designing and implementing effective industrial automation systems presents many technical challenges. Because the lifecycle of an industrial plant often far exceeds the lifecycle of the technology that runs the plant, the management and maintenance costs of the technology are often difficult to manage. In the following example, resource abstraction allows SDIS deployment to adapt to the dynamic configuration (and reconfiguration) of software and hardware resources in an industrial system. Such resource abstraction provides the flexibility to update configurations without causing the industrial system to cease operation; it also provides the flexibility to update industrial systems with improved capabilities over time.

[0084] The use of an open architecture and abstract links between software and hardware in the currently disclosed SDIS approach provides these and other technical benefits while allowing vendors to focus on the capabilities and implementation of specific vendor applications. The disclosed open architecture also promotes innovation, reduces the cost of hardware replacement, and eliminates the risk of hardware obsolescence. The disclosed open architecture enables security to be implemented as an inherent part of the SDIS, such as through the use of hardware trust roots, signed applications, and comprehensive security management. Such a configuration enables a simplified control system with inherent security and the ability to easily integrate capabilities over time. These technical improvements, combined with the characteristics of open architecture and standard implementation, enable the rapid integration of industrial controls in the SDIS.

[0085] Several existing initiatives, such as the Open Process Automation Forum of the Open Group, have begun developing standards-based, open, interoperable process control architecture features for industrial automation, targeting industries such as food and beverage, mining and metals, oil and gas, petrochemical, pharmaceutical, pulp and paper, and utilities. The current configuration and functionality of the SDIS and accompanying subsystems and technologies can be integrated into industrial automation and system deployment efforts using this standard or similar approaches. Furthermore, the current configuration and functionality of the SDIS and accompanying subsystems can be used in these and other industries. Therefore, variations and modifications to the following implementation will be readily apparent.

[0086] Figure 1A A first example configuration of the SDIS operational architecture is depicted. As shown, a control message bus 112 is used to connect the various components of the architecture, including an operational tool 120, a control server (CS) node 130A, an edge control node (ECN) system 150, an intelligent I / O controller system 165, a base I / O controller system 160, a gateway system 170, and a control station 115. Various field devices (151, 161, 166, 171) are connected to the various systems (150, 160, 165, 170). Some example use cases and configurations of this operational architecture are discussed further below.

[0087] In an example, the operation tools 120 may include the following aspects: program development tools, history tools, human machine interface (HMI) development, control and operation tools. Figure 2A ) to implement various aspects of the operating tool 120.

[0088] In an example, the control server node 130A may include aspects of various virtual machines 131A that are coordinated through a hypervisor layer 132A and operate using features of a host operating system 133A and computer hardware architecture 134A. The control server node 130A may be used to implement various aspects of orchestration 135A, involving both machine orchestration and operational application orchestration. Figure 2A Further detailed discussion of the control server node 130A is provided.

[0089] In an example, the ECN system 150 may include aspects of an orchestration (e.g., an orchestration implementation) from an ECN I / O controller (e.g., nodes 150A, 150B) operating on specific hardware (e.g., x86 or ARM hardware implementation). Figure 2B Further detailed examples of the ECN system 150 and its role in orchestrating various connected devices (eg, field devices 151A, 151B) are provided.

[0090] In an example, the intelligent I / O system 165 may include various configurable aspects of industrial control from an intelligent I / O controller (e.g., controllers 165A, 165B) and an accompanying operating system for controlling or accessing various devices (e.g., field devices 166A, 166B). Also in an example, the basic I / O system 160 may include various operational aspects of industrial control from a basic I / O controller (e.g., controllers 160A, 161B) and an accompanying operating system for controlling or accessing various devices (e.g., field devices 161A, 161B).

[0091] In an example, the gateway system 170 may include various configurable aspects for connecting to other device networks or deployments from a gateway (e.g., gateways 170A, 170B) for control or access of various devices (e.g., field devices 171A, 171B). In the various devices, the roles of sensor ("S") and actuator ("A") components are labeled throughout the field devices (e.g., on field devices 151A, 151B, 161A, 161B, 166A, 166B, 171A, 171B). It will be understood that additional numbers and types of devices and components may also be coupled to the various systems 150, 160, 165, 170.

[0092] Figure 1AThe operational architecture depicted is configured to achieve many of the same attributes seen in traditional enterprise architectures, such as HW / SW modularity, SW portability, interoperability, application extensibility, and compute scalability. Furthermore, the new infrastructure framework components introduced in this architecture can be deployed, most notably in the implementation of CS and ECN systems, to support the centralized and decentralized concepts of the SDIS technology discussed in this article.

[0093] For example, the use of ECN I / O controllers (e.g., in ECN nodes 150A, 150B) is architecturally significant compared to the current DCS (distributed control systems) and PLC (programmable logic controller) control systems that have evolved over the past fifty years. Any architectural advancements within this mission-critical portion of the ANSI / ISA-95 automation interface stack must adhere to the stringent and resilient requirements of process control. Leveraging the SDIS architecture described herein, ECN systems can not only maintain these stringent operational requirements but also remain open and interoperable, while allowing industrial users to confidently, reliably, securely, and quickly introduce or refresh these systems through ongoing technological advancements. This SDIS architecture enables broader ecosystem participation, innovation, and product customization across the entire operations and control stack. For example, control decomposition can be provided for ECN systems to serve as fundamental control system building blocks, enabling amplified control function customization and increased processing flexibility for a variety of use cases.

[0094] Figure 1B A second example configuration of the SDIS operational architecture is depicted. Figure 1A In a similar manner as shown, Figure 1B The configuration diagram shows a control message bus 112 that is used to connect the various components of the operating architecture, including cloud components (real-time advanced computing system 130B working as a control server, and cloud computing service 180), edge components (edge ​​ecosystem 190 consisting of edge computing nodes 191A, 191B, 191C, first edge computing platform 193 and second edge computing platform 195), and control station 115. Various field devices 192, 194 with sensors and actuators are connected to the corresponding edge computing nodes (in edge ecosystem 190 and edge computing platforms 193, 195). The operational goals and characteristics discussed above also apply to Figure 1B Configuration.

[0095] As Figure 1A A further extension of the SDIS operational architecture introduced in Figure 1B The configuration diagram shows a scenario where the operations of the controllers and servers across various cloud and edge components are virtualized through corresponding virtual machines, deployed with corresponding containers, deployed with corresponding applications, or any combination thereof. Figure 1B The SDIS operating architecture allows for reconfigurable and flexible deployment on a variety of hardware setups, including both ARM and x86 hardware architectures. Figure 3A A further breakthrough in real-time advanced computing systems 130B is depicted in Figure 3B and Figure 3C Further breakthroughs in cloud computing service nodes 180 and edge computing nodes 193 are discussed separately.

[0096] Another aspect of the SDIS architecture can involve the use of real-time communications. A control message bus 112 hosted on the service bus fabric 110 can be used to enable internetworking convergence at multiple levels. For example, the control message bus 112 can enable the use of time-sensitive Ethernet transport, such as through the Ethernet-based Time Sensitive Networking (TSN) open standard (e.g., IEEE 802.1TSN Task Group). Furthermore, the use of the control message bus 112 can enable higher performance and scale at the cloud server rack level and across large networked edge nodes or edge node architectures.

[0097] In the SDIS architecture, real-time services can operate on top of a real-time physical transport (such as Ethernet TSN) via a control message bus 112. The control message bus 112 can be adapted (e.g., by using Open Platform Communications Unified Architecture (QPC-UA), Object Management Group Data Distribution Service (DDS), OpenDXL, Open Connectivity Foundation (OCF), or similar standards) to address the heterogeneity of existing middleware or communication stacks in IoT settings to enable seamless device-to-device connectivity to address emerging implementations of IoT deployments.

[0098] In an example, orchestration management for the SDIS architecture may be implemented through a control server (CS) design. Figure 2A The diagram shows the SDIS operating architecture (eg, Figure 1A Specifically, Figure 2A A further illustration of the CS node 130A and its components virtual machines 131A, hypervisor 132A, host operating system 133A, and hardware architecture 134A is provided; as shown, the CS node 130A is shown as a single node, but may comprise two or more nodes with many virtual machines distributed across these nodes.

[0099] In an example, CS node 130A may include orchestration 135A, which is facilitated by machine and operational application orchestration. Machine orchestration may be defined using a machine library 136 (such as a database used to implement platform management); operational application orchestration may be defined using a control function library 142 and an operational application library 144. For example, control standard design 141 and integrated (and secure) application development process 143 may be used to define libraries 142 and 144.

[0100] In an example, the CS node 130A is designed to host ISA level L1-L3 applications in a virtualized environment. This can be achieved by running virtual machines (VMs) 131A on top of a hypervisor 132A, where each VM encapsulates a Future Airborne Capability Environment (FACE)-compatible stack and applications, or non-FACE applications such as human-machine interfaces (HMIs), historians, operational tools, etc. In an example, a FACE-compatible VM can provide the entire FACE stack (operating system, FACE segment, and one or more portable components) encapsulated in the VM. Encapsulation means that each VM can isolate its own virtual resources (compute, storage, memory, virtual network, QoS, security policy, etc.) from the host and other VMs through the hypervisor 132A, even though each VM may be running a different operating system, such as Linux, VxWorks, or Windows.

[0101] To maximize the benefits of virtualization and robustness, related groups of portable components can be grouped in FACE-compliant VMs and multiple FACE-compliant VMs can be used. This approach can distribute workloads across CS hardware and isolate resources specific to that group of components (such as networking), while still allowing applications to communicate with other virtualized physical devices (such as ECN) over the network. Distributing FACE portable components across VMs increases security by isolating unrelated components from each other, provides robustness to failures, allows independent updates of functionality, and simplifies integration, allowing independent vendors to contribute fully functional VMs to the system.

[0102] In a further example, layer 2 components can be separated from layer 3 components in separate VMs (or VM groups) to provide isolation between the layers and allow different network connections, security controls, and monitoring to be implemented between the layers. Grouping portable components can also provide benefits for integration, allowing multiple vendor solutions to be easily combined to run multiple virtual machines and configure networks between them. Also in a further example, additional operating systems such as Windows, Linux, and other Intel architecture-compatible operating systems (e.g., VxWorks real-time operating system) can each be deployed as a virtual machine. Other configurations of the currently disclosed VMs within the CS node 130A can also achieve other technical benefits.

[0103] In an example, a cloud infrastructure platform, such as a real-time advanced computing system adapted to use open source standards and implementations such as Linux, KVM, OpenStack, and Ceph, can be utilized in the CS node 130A. For example, the cloud infrastructure platform can be adapted to address key infrastructure requirements such as high availability of the platform and workloads, continuous 24 / 7 operation, determinism / latency, high performance, real-time virtualization, scalability, upgradeability, and security. The cloud infrastructure platform can also be adapted to meet software-defined key infrastructure requirements specific to industrial automation.

[0104] Figure 2B The diagram shows the SDIS operating architecture (eg, Figure 1A

[0026] An example configuration of a distributed edge control node (ECN) subsystem within an operational architecture discussed in the ISA-95 documentation. In the example, the ECN nodes 150A, 150B reside in ISA-95 Level 1 / Level 2 and are positioned as basic foundational HW / SW building blocks.

[0105] In an example, ECN nodes 150A, 150B support a single input or output to a single fieldbus device via a sensor, actuator, or intelligent device (e.g., located external to the ECN cabinet). The ECN device architecture can be extended through an ECN cabinet or rack system, thereby addressing the cabling, upgrade, and fault tolerance limitations of existing proprietary DCS systems. The ECN cabinet or rack system extends the openness and flexibility of the distributed control system. In an example, the ECN architecture operates within a standard POSIX OS, where a FACE-compliant stack is implemented as a segment or group of software modules. Various methods for deploying these software modules are referenced in the following examples.

[0106] The ECN nodes 150A, 150B may support various software-defined machines for various aspects of orchestration and services (such as those described below for Figure 6In an example, ECN nodes 150A, 150B may be integrated with various hardware security features and trusted execution environments, such as Software Guard Extensions (SGX), Dynamic Application Loader (DAL), Secure VMM Environment and Trusted Computing-Standard Trusted Platform Module (TPM). In a further example, the system can be protected by hardware encryption engines such as Secure boot is enabled by converging and protected key material accessed from the Converged Security and Manageability Engine (CSME) and Platform Trust Technology (PTT). In addition, special hardware instructions for AES encryption and SHA calculations can be used to make encryption functions more secure. Other forms of security, such as Enhanced Privacy ID (EPID), which can be adopted throughout the industry as the preferred device identity key, can be enabled through automated device registration (e.g., Intel Secure Device Activation (SDQ)) technology for secure, zero-touch onboarding of devices. In further examples, ECN nodes 150A, 150B and other subsystems of the SDIS architecture can interoperate with these or other secure approaches.

[0107] Figure 3A The diagram shows that Figure 1B More detailed configuration of the real-time advanced computing system 130B deployed within the SDIS operating architecture. Specifically, Figure 3A The configuration diagram illustrates the operation of various virtual machines 131B operating on a hypervisor layer 132B, which may include virtual machines, containers, and applications of different deployment types. When the VMs, hypervisors, and operating systems are executed on a hardware architecture 134B (e.g., a commercial off-the-shelf (COTS) x86 architecture), a host operating system 133B may be used to control the hypervisor layer 132B. Various aspects of real-time orchestration 135B may be integrated into all levels of computing system operation. Thus, an x86 computing system may be suitable for coordinating any of the cloud-based or server-based SDIS functions or operations discussed herein. Other aspects of the functionality or hardware configuration discussed for CS node 130A may also be applicable to computing system 130B.

[0108] Figure 3B and Figure 3C The diagrams show that Figure 1B A more detailed configuration of the cloud computing 180 and edge computing 193 subsystems deployed within the SDIS operational architecture. Figure 3A In a similar manner as described in Figure 3B and Figure 3CThe series of virtual machines 181, 196, hypervisor layers 182, 197, host operating systems 183, 198, and COTS x86 hardware architectures 184, 199 depicted in FIG can be adapted to implement the respective systems 180, 193. Applications and containers can be used to coordinate cloud-based and edge-based functions under the control of real-time orchestration. Other aspects of the functions or hardware configurations discussed for ECN node 150 can also be applied to edge computing node 193. Edge computing node 193 can implement control functions to control field devices.

[0109] The systems and techniques described herein can integrate "mobile edge computing" or "multi-access edge computing" (MEC) concepts, which access one or more types of radio access networks (RANs) to allow for an increase in the speed of content, services, and applications. MEC allows base stations to act as intelligent service hubs capable of delivering highly personalized services in edge networks. MEC provides a close, fast, and flexible solution for a variety of mobile devices, including those used in next-generation SDIS operating environments. As an example, the paper "Mobile-Edge Computing, A key technology towards 5G," published by the European Telecommunications Standards Institute (ETSI) as ETSI White Paper No. 11, authored by Yun Chao Hu et al., with an International Standard Book Number 979-10-92620-08-5 and available at http: / / www.etsi.org / news-events / news / 1009-2015-Q9-news-new-white-paper-etsi-s-mobile-edge-comiiting-initiative-explained, describes an MEC approach, which is incorporated herein in its entirety. It will be understood that other aspects of 5G / next generation wireless networks, software-defined networks, and network function virtualization may be used with the present SIDS operating architecture.

[0110] Figure 4 The diagram illustrates an example configuration 400 of a real-time service bus (e.g., a configuration of the control message bus 112) for use within the SDIS operating architecture. As discussed herein, this configuration allows for support of various processing control nodes. For example, the control message bus 112 can be used to connect various control processing nodes 410 (including various hardware and software implementations on nodes 410A, 410B, 410C) and cloud-based services or control servers 130A with various edge devices 420 (e.g., I / O controllers 150, 160, 165 or edge computing nodes 191, 193, 195).

[0111] In an example, the control message bus 112 can be implemented to support packet-level, deterministic control networks with rate-monotonic control requirements. These features are typically provided by proprietary distributed control systems (DCS), supervisory control and data acquisition (SCADA), or programmable logic controller (PLC) components. Most of these systems are designed with design parameters that limit the number of nodes and data elements, with little ability to dynamically manage the quantity and quality of data across the typically closed and isolated networks within a facility. Over the lifecycle of these systems, the willingness to implement emerging new use cases is severely limited by the potential inflexibility and limited scalability of the expensive control system infrastructure.

[0112] Leveraging previous approaches, both open source and open standards-based service bus middleware options have matured to the point where the mission-critical ecosystem of solution providers has come to view these technologies as “best-in-class” capabilities for building scalable, highly redundant, fault-tolerant, real-time systems at a fraction of the historical cost. This has spurred the realization of new use cases for both discrete and continuous processing where commodity-grade hardware and open source, standards-based software have converged to enable real-time computing approaches while maintaining design principles based on service-oriented architectures.

[0113] In one example, control message bus technology can be further extended by enabling time-sensitive networking (TSN) and time-coordinated computing (TCC) between and within platform nodes of the network, thereby enabling real-time computing at the hardware level. Both proprietary and open standards-based solutions can be integrated with enhanced capabilities to support commodity hardware, including leveraging industry standards provided by the OPC-UA (OPC Unified Architecture) and DOS (Data Distribution Service) groups, as well as proprietary implementations such as the SERCOS standard, where hard real-time requirements for discrete motion control are mandatory in robotics and machine control applications.

[0114] In an example, the control message bus and the entire SDIS architecture can also be integrated with Industrial Internet Consortium (IIC) features. These can include various developed and tested standards for industrial use of TSN, which can enhance the performance and QoS of DDS-based solutions and OPC-UA-based solutions by significantly reducing packet-level latency and jitter. In addition, various aspects of the Object Management Group (QMG) and OPC Foundation standards can be positioned to support the integration of increased OPC-UA and DDS implementation models that leverage the information modeling of OPC-UA and the QoS and performance capabilities of DDS in architectural design. New use cases can include analytics and autonomous capabilities.

[0115] In an example, the SDIS architecture can be integrated with the use of software-defined networking (SDN) features. SDN is the evolution toward software-programmable networks that separate the control plane from the data plane to make networks and network functions more flexible, agile, scalable, and less dependent on networking devices, vendors, and service providers. Two key use cases for SDN related to SDIS include service function chaining, which allows for dynamic insertion of intrusion detection / prevention functions, and dynamic reconfiguration for responding to events such as large-scale outages (such as regional maintenance, natural disasters, etc.). Additionally, the SDIS architecture can be integrated with an SDN controller to control virtual switches using networking protocols such as the Open vSwitch Database Management Protocol (QVSDB). Other use cases for SDN capabilities may involve dynamic network configuration, monitoring, and abstraction of network functions in virtualized and dynamic systems.

[0116] Figure 5A The diagram illustrates a first network configuration 500 for an example deployment of the SDIS subsystem. The first network configuration 500 illustrates a scaled-down, small-scale deployment option that combines controller storage and compute functionality on a redundant pair of hosts (nodes 510A, 510B). In this configuration, the controller functionality (for controlling applications or implementations) is in an active / standby state on nodes 510A, 510B, while the compute functionality (for all remaining processes) is in an active / active state, meaning that VMs can be deployed to perform compute functionality on either host.

[0117] For example, LVM / iSCSI can be used as a volume backend for replication across compute nodes, while each node also has local disks for temporary storage. Processor bandwidth and memory can also be reserved for controller functions. When less processing and redundancy are required, this two-node solution can provide a lower cost and smaller footprint solution.

[0118] Figure 5B The diagram illustrates a second network configuration for deployment of the SDIS subsystem. The second network configuration 550 can provide high capacity, scalability, and performance for dedicated storage nodes. Compared to the first network configuration 500, the second network configuration 550 allows controller, storage, and compute functions to be deployed on separate physical hosts, thereby allowing storage and compute capabilities to be scaled independently of each other.

[0119] In an example, the second network configuration can be provided by a configuration of up to eight storage nodes (nodes 530A-530N) and eight disks per storage node in a high-availability (e.g., Ceph) cluster (e.g., coordinated by controller nodes 520A, 520B), with the high-availability cluster providing image, volume, and object storage for the compute nodes. For example, up to 100 compute nodes (e.g., node 540) can be supported, each with its own local temporary storage for use by VMs. As will be appreciated, a variety of other network configurations can be implemented using the present SDIS architecture.

[0120] The SDIS architecture, along with its accompanying data flows, orchestration, and other features expanded upon below, can also leverage aspects of machine learning, cognitive computing, and artificial intelligence. For example, the SDIS architecture can be integrated with a reference platform that has a foundation in hardware-based security, interoperable services, and open source projects, including the use of big data analytics and machine learning for network security. The SDIS architecture can leverage immutable hardware elements to demonstrate device trust and characterize network traffic behavior based on filters enhanced by machine learning to separate bad traffic from good traffic.

[0121] The various components of the SDIS architecture can be integrated with a rich set of security capabilities to enable interoperable and secure industrial systems in real-world industrial settings. For example, such security capabilities can include hardware-based roots of trust, trusted execution environments, protected device identities, virtualization capabilities, and cryptographic services on which a robust real-time security architecture can be built. The configuration and functionality of such components in a functional SDIS architecture deployment are further discussed in the following sections.

[0122] Overview of the data model

[0123] In this example, the SDIS architecture can be further integrated with various data models for managing data from sensors, actuators, and other deployed components. A time series stream of numerical data is of limited use if one does not know how or where the data was generated, what measurements are being collected, or other characteristics. Data models can be used to provide context for this information across many domains, even in extreme cases where users wish to obfuscate the identity of the data.

[0124] Data models can be defined to provide a representation of data structures. Data models can also be defined to allow different stakeholders to define multiple objects and how these objects interact or relate to each other. For example, semantic data models can be used across multiple domains and assist in processing and storage in various information systems within the SDIS architecture.

[0125] In the example, a semantic data model can define aspects of any combination of the following components:

[0126] Metadata: (ie, information describing what the data is about).

[0127] For example, a data stream or data point may have metadata containing a name, such as "temperature." Another piece of metadata may be a location, "second floor, pole J2," to indicate the source of the data. Furthermore, such metadata can be flexible and extensible.

[0128] Taxonomy: In taxonomy, data can describe the categories and relationships between data points. Taxonomy can include information about the analytics performed on a piece of data and how that data relates to other data or devices in a specific site. Tag libraries can be defined for the system to ensure interoperability and support for multiple devices.

[0129] Object structure: The object structure can be used to describe what metadata and classifications an object may and should have.

[0130] Data flow: A data flow can describe data transformations and data flows, and such a data flow can be abstract or physical. In a further example, a data flow can rely on a standard definition or approach such as REST.

[0131] Data storage: The storage and use of data in a specific data storage configuration may affect the data model and the performance of data producers and consumers.

[0132] As expanded upon in the examples below, the SDIS architecture can provide a common data model to address heterogeneity in data propagation across applications and machines. As also discussed below, dynamic data models can be leveraged within the SDIS architecture to provide an abstract representation of data structures and further allow different stakeholders to define flexible data objects, their representations, and their interrelationships. These and other semantic data models can be essential for processing and storage in many information system deployments.

[0133] Dynamic Data Model

[0134] As mentioned above, a data model can be a fundamental component used in IoT deployments (such as SDIS implementations). A data model is an abstraction of data and relationships across different structures and flows. Based on the implementation, the data model can be implemented through simple on-the-fly tagging (e.g., as used in Project Haystack, an open source initiative for developing naming conventions and taxonomies for building device and operational data) or through extensive definitions of structures / classes and data flows (e.g., such definitions are typically established during the design phase and before system development). The data model is important in many systems because it provides a mechanism for developers, designers, architects, and deployment technicians to describe and find data sources.

[0135] Most data models involve time and effort (and multiple iterations) to produce their definitions. Furthermore, most data models are static and require considerable modification and iteration to add new tags, components, or connections to existing models, often making them backwards incompatible. This prevents extensive changes to data models during deployment, which are used to describe tags used by data applications (such as in data visualization, analytics, or provisioning tool applications).

[0136] While existing solutions offer limited flexibility in defining data models, they are not dynamic. For example, while a data designer can define a device with many features, some of which are optional, certain features may not be changed during runtime. This makes the creation, modification, and maintenance of data models very complex, especially in industrial environments.

[0137] By definition, data model creation tends to be static. This means that once a data model is defined and implemented, any changes to the structure (e.g., the structure of the data model, not the data values) typically require the development and deployment of a new version of the code for that data model. However, this does not address scenarios where applications, devices, or sensors have dynamic feature sets. An example would be a device that acts as a sensor ensemble, representing itself as different output sensors based on battery and computational availability. Such a device might contain multiple physical sensors (proximity, contact, light, Bluetooth, temperature, etc.) and could report on the occupancy of a room. However, if for some reason the device needs to conserve power or encounters a faulty module, it can revert to a subset of its components. In this context, the concept of a dynamic data model (and thus dynamic metadata and features) is extremely valuable, even with complex representations such as probabilistic estimates of tags rather than binary on / off states. The concept of a dynamic data model can also provide valuable input for predicting when orchestration should (or should) be deployed in IoT environments.

[0138] The following techniques enable the creation and deployment of dynamic data models to address these and other technical considerations in SDIS architectures and similar settings. In a dynamic data model, the data designer can identify a set of required fields, such as name, unit, type, etc. Additionally, the data designer can leave the definition open for nodes and modules to add metadata (of any type and amount), or the definition can restrict what can be added and by whom.

[0139] In this example, when a node observes certain behavior in a data stream, it can query the sensor's metadata for expanded rules. For example, a module can use analytics to generate predictive maintenance outputs. As part of the computation, if the model applies unsupervised learning, it can detect that a particular sensory data stream is becoming increasingly important over time and should be added to the feature set. Feature sets are crucial for analytics computations and, depending on the application, can also lead to real-time requirements.

[0140] For example, adding this flag to metadata would allow TSN switches to prioritize network traffic to support learning algorithms. When dynamic metadata is added, data is assigned a validity period or refresh period. This ensures that the data continues to support a specific feature, otherwise the feature may no longer be valid and can be updated accordingly. Alternatively, a discard mechanism could be implemented that allows the system to discard multiple pieces of metadata when they are no longer valid or needed.

[0141] In the example implementation, each data flow is assigned to a data flow manager. The data flow manager can be the device generating the data, or a virtual implementation located in the fog or cloud. The data flow manager carries the policy for dynamic metadata. When another module or node in the system needs to act on the dynamic data model, it contacts the flow manager. The flow manager can then provide it with a token with an equivalent key, allowing the data flow manager to add and update metadata as the data flow manager observes and analyzes the traffic.

[0142] In a further example implementation, the system can add provenance metadata. Provenance provides a trace of ownership and modifications that nodes can use to understand the history of the stream or just parts of the metadata.

[0143] Figure 6The diagram illustrates a protocol for establishing a dynamic data model according to an example. In this example, a sensor 610 generates streaming data in a data stream 630. The data stream is acquired and processed by multiple nodes (e.g., servers 640 and 650 monitoring the data stream 630). The sensor data generated in the data stream 630 is marked to indicate that the sensor 610 supports dynamic data modeling. Then, a data model manager (e.g., operating on the server 640) acquires and processes the sensor data.

[0144] The data model manager can reside anywhere in the network as long as it is accessible to the sensor 610 and other modules and devices that are allowed to modify the data model of the subject sensor. For example, as data flows upstream, a device (such as server 650) can run a set of analytics to determine whether the data stream should be considered in the feature set used for the algorithm. Due to a change in the system, the device determines that the data stream generated by the sensor 610 is now valuable for controlling the process of the robot 620. The device then sends a request with its credentials to the data model manager to request to modify the sensor data stream or add a flag to the sensor data stream that will indicate relevance to the robot arm. The data model manager determines whether the algorithm in question has the authority to request such a change.

[0145] Using predefined policies, the data model manager sends commands to the device requesting the implementation of the modifications requested by the algorithm. These changes will take effect based on the policy or can be part of the request. The newly added metadata to the data model can have further consequences, such as affecting other factors such as connectivity QoS (for example, involving TSN-enabled communication). Similarly, if for some reason the algorithm or application is no longer interested in the sensor data, the data model can be modified to omit the tag in question. The request can include the complete removal of the tag or even a temporary suspension before further data is analyzed.

[0146] Figure 7 A flowchart 700 illustrates an example process for dynamically updating a data model generated and utilized in an SDIS operational architecture. As depicted, the following flowchart includes a number of high-level operations that may be performed by one or more systems or subsystems (e.g., Figure 6 However, in the SDIS architecture, the following operations can be adapted between various computing or controller nodes for use with various connected sensors and actuators.

[0147] In flowchart 700, operations include monitoring a data stream provided by a sensor or actuator in a controlled system (operation 710). This monitoring can be provided continuously in the data stream, by sampling data from a data source, or by any number of monitoring methods. Based on this monitoring, one or more patterns can be detected from the data stream (operation 720). For example, data values, data value trends, data value confidence or probability, or other combinations of values ​​indicated by one or more data value types and sources can be analyzed for one or more patterns. In examples, machine learning, artificial intelligence, or rules can be employed for this pattern detection.

[0148] One or more detected patterns can be used to identify data model changes (operation 730). For example, certain combinations of data values ​​can be used to trigger the addition of an aggregate data value type to be added to the data model; similarly, for example, a trend or confidence level in a data value can cause a data value type to be deleted or changed. The identified data model changes can then be incorporated into the data model (operation 740) and deployed for use in various system components. Accordingly, subsequent system operations (including system commands and workflows) can be executed in the system deployment based on the data model changes (operation 750). Other types of data analysis and system operation adaptations may also occur as a result of the data model changes.

[0149] As an extension of the dynamic data model changes discussed above, the presence of labels in a data stream can also be used to indicate the confidence or data type of a data value. For example, suppose a feature selection component is used to determine the relevance of a particular data stream to an analytical function. The feature selection component can determine a relevance score for that data stream. As a further example, the feature selection component can generate a relevance score of 0.8 for a particular information field in the data stream. Such a relevance score can be used as a confidence level defined in metadata, which will be added to the data model.

[0150] In a similar manner, the same data stream may have a very low relevance score (0.4) for another information field such as occupancy. Another device or algorithm may query the device for its metadata with the filter set to high confidence. As a result, the device will return metadata associated with a relevance score of 0.8, but will omit metadata with a relevance score of 0.4.

[0151] In such an example, not only is the data model dynamic, but the relevance score used to evaluate a particular data field can also be dynamic and can be recalculated periodically or based on events. In yet another example, the relevance score is not defined as a single value, but is represented by a vector having a set of conditions associated with multiple values. An example of a vector can be as follows:

[0152] Tags: "algorithm": "occupancy"

[0153] Confidence vector [0.7, 0.3, 0.8]

[0154] Context vector [“7:00am-5:00pm”, “5:01pm-8:00pm”, “8:01pm-6:59am”]

[0155] In this example, there are three confidence levels with three associated contexts. The context can be as simple as the time of the data, or more complex to provide event-based expressions.

[0156] Continuing with the previous example, consider a scenario where light sensors are used to determine room occupancy in a smart building deployment. The sensor value can provide an accurate indication of occupancy status during normal office hours. However, after cleaning staff hours, when multiple lighting fixtures rapidly turn on and off, this is not normal, and the sensor value can be thrown off as unusual activity by the cleaning staff. However, after 8:00 pm, when cleaning staff are typically absent, the presence of lighting can again serve as an accurate indicator of occupancy.

[0157] A dynamic data model can allow for the useful addition and removal of tags based on the generated context and data (or attributes of the data). This can allow SDIS deployments to add traffic prioritization, policies, and even routing decisions based on those dynamic tags without having to recreate a new data model to add or remove those extensions.

[0158] From a developer and application perspective, to support dynamic data models, each device must support, modify, and add appropriate queries to the device's data model. Devices must also support queries that return data models based on a set of criteria. Consequently, device interfaces for modifying, adding, and returning data models are used for runtime modification of data models.

[0159] In a further example, data models can also be synchronized across multiple devices and nodes. For example, an algorithm may determine, based on data from one portion of a building in a deployment, that a certain sensor typically placed in a conference room is now highly relevant to an occupancy algorithm. If this is the case, the sensor's data model can be modified to reflect this new finding. Additionally, other sensors can be asked to execute a piece of code to determine if they exhibit similar behavior. If feature extraction shows similar behavior, the data model can also be expanded. As a result, application developers and even system integrators can allow modifications to occur even if they have not been validated on all sensors. Such a capability can be used as a compromise between accuracy / validation and the assumption that similar sensor data is valuable and should be routed and managed accordingly.

[0160] Tags or metadata that are dynamically added to the data model can also be allowed to age (e.g., decay) and be removed or demoted (by reducing the relevance score) unless the device or algorithm continues to verify the need and relevance of such metadata. Such aging can allow for natural pruning of metadata that has become obsolete even if the developer did not recognize it. In basic implementations, aging can be strictly based on time; however, other implementations can include advanced concepts, including aging based on lack of usage. Similarly, if the device or algorithm requested to be added does not consume data from the sensor, the metadata can be aged or archived (e.g., continue to be available, but not given any priority). However, periodic use of metadata (notwithstanding queries or other QoS decisions made by the system) can keep the interest in the metadata current.

[0161] Figure 8 A flowchart 800 illustrates an example method for maintaining various aspects of a dynamic data model in an SDIS operational architecture. In an example, the method may include: identifying optional preconditions for one or more conditions for data model evaluation (operation 810); acquiring data from one or more sensors via the data stream for data provided in a data stream according to the data model (operation 820); identifying one or more thresholds for data model modification (operation 830); and evaluating the data from the sensor(s) for data model modification using pattern(s) or rule(s) and the identified threshold(s) (operation 840).

[0162] The method may also include defining feature additions, changes, or deletions for data model modification (operation 850); requesting approval of the data model modification from the data model manager (operation 860); receiving approval of the data model modification from the data model manager and processing the approval of the data model modification (operation 870); incorporating the data model modification into the data model for one or more sensors or data streams (operation 880); and implementing changes to data processing in the system architecture based on the data model modification (operation 890).

[0163] Any of these dynamic data model operations may be expanded based on other examples, scenarios, or conditions discussed above. Additionally, additional aspects of maintaining and utilizing dynamic data models may be combined with the functional orchestration or other management features of the SDIS architecture disclosed herein.

[0164] Overview of Functional Orchestration

[0165] Figure 9The diagram illustrates an example of a set of orchestration operations 900 dynamically established using the Composable Application System Layer (CSL) in the SDIS operational architecture. The CSL can be utilized to implement secure design and orchestration of control functions and applications to support industrial operations.

[0166] In an example, CSL maintains a library 980 of function blocks 990, each of which identifies control loop logic and application components. Each function block can interoperate with other function blocks. Function blocks can have multiple implementations, making them portable, allowing them to run on various platform architectures and take advantage of special features (e.g., hardware accelerators) when available. In an example, CSL provides control functionality for a cluster of edge nodes (e.g., ECN); in a further example, CSL provides control for VMs in control servers or other computing points in the SDIS operating architecture.

[0167] In an example, a process engineer (or other operator) defines control flows and applications by combining and configuring existing function blocks 990 from a library 980. These function blocks 990 can represent application logic or control loops (e.g., control loop 970, data storage, analytics modules, data acquisition or execution modules, etc.), control modules, or any other computing elements. Because these function blocks 990 are reusable and interoperable, new code only needs to be written when a new function block is needed. In a further example, such function blocks can be used to implement end-to-end logic, including control flows or end-to-end applications using a graphical drag-and-drop environment.

[0168] Starting with the application design, CSL generates an orchestration plan 940 that specifies the required functional blocks and the requirements for the computational points used to execute these functional blocks. As discussed in the following sections, orchestration 920 can include a process for mapping orchestration plan 940 to available computational and communication resources. Orchestration 920 can be further adapted based on control standard design 910 (e.g., to make the resulting orchestration conform to various control laws, standards, or requirements).

[0169] In the example, the CSL maintains a map 930 of compute and control resources across the SDIS network. Map 930 includes the topology of various compute points, from virtual machines in the data center to control points and the attached sensors and actuators. Map 930 also includes the hardware capabilities and dynamic characteristics of the control points. Map 930 is regularly updated, allowing the system to continuously adapt to component failures. Orchestration 920 and control loop 970 communicate using monitoring logic 950 and function deployment 960. Monitoring logic 950 outputs information from field devices or control loop 970, which serves as input to map 930. Function deployment 960 serves as input or state setting for control loop 970.

[0170] When an operator deploys a new application definition (e.g., orchestration 920 receives output from control standard design 910), orchestration 920 determines how to best fit function block 990 into the set of available resources in mapping 930 and deploys the underlying software components that implement function block 990. Deployment of an end-to-end application can include, for example, creating virtual machines within a server as needed, injecting code into control loops (e.g., control loop 970), and creating communication paths between components. Orchestration 920 can also be dynamic to allow migration of function blocks in the event of a computing resource failure without requiring a system-wide restart. Additionally, updates to component implementations can be pushed, causing the code to be updated as needed.

[0171] The CSL may also incorporate security and privacy features, such as to establish trust with participating devices (including edge nodes or control servers). In a further example, the CSL may be integrated with key management for onboarding new devices and for decommissioning obsolete devices. The CSL may pass keys to a functional block 960 to enable secure communication with other functional blocks 960. The CSL may also deliver secure telemetry and control, integrity and isolated execution of deployed code, and communication integrity between functional blocks 990.

[0172] Additional examples of orchestration functionality developed within the orchestration architecture of the SDIS are discussed further in the following sections.

[0173] Orchestration of distributed mission-critical workloads

[0174] Today's orchestration technologies are primarily implemented using functions, applications, virtual machines, or containers. However, today's control strategy implementations often fail to manage the inherent dependencies between distributed applications within low-latency, high-frequency, mission-critical timeframes. Historically, dynamic orchestration has not been applied to embedded systems due to technical limitations in managing application dependencies at runtime.

[0175] The following techniques address the orchestration of distributed workloads defining real-time, mission-critical control strategies for industrial systems. The orchestrated control strategies can operate within the new distributed control system (DCS) designs currently being defined and can be applied to discrete, continuous, and batch manufacturing operations. For these systems, real-time, mission-critical control applications can be constructed following the IEC 61499 standard and represented as a combination of multiple scheduled and coordinated synchronous or asynchronous, event-driven building block applications. Where application functionality is unique, the building blocks can be executed consistently in a specific order and at a frequency within defined system latency boundaries.

[0176] In contrast, many existing embedded applications typically run on dedicated, fixed-purpose hardware. Traditional application orchestration does not consider the application processing dependencies on other application building blocks that make up the complete control strategy, where the total compute, memory, storage, and scheduling need to execute together for the mission-critical control strategy to execute without error.

[0177] In an example, features of the SDIS architecture can be adapted to support the overall orchestration and management of multiple dependent applications (functional blocks) executed across a distributed resource pool, so as to enable orchestration at the embedded control policy level within a distributed system configuration. This provides control policy orchestration capabilities to the operational technology environment while improving overall system performance at an expected reduced total cost. For example, an example orchestration approach can incorporate dynamic network discovery, resource simulation prior to any orchestration action, and simulation combined with global resource optimization and forecasting as part of the orchestrator's rule set decision tree.

[0178] The distributed resource pool may include applications spanning the following ranges: (a) a single application running on a single native device, with a second, redundant application available on a nearby native device; (b) multiple coordinated applications running on multiple native devices; (c) multiple coordinated applications running on a single virtual machine running on a single embedded device or server; (d) multiple coordinated applications running across multiple virtual machines, each running on a dedicated embedded device or server; (e) multiple coordinated applications across multiple containers contained in a virtual machine running on a dedicated embedded device or server; or (f) multiple coordinated applications across multiple containers running on multiple embedded devices or servers. Any mixture of these application scenarios may also be applicable.

[0179] In an example, orchestration may include measurement of resources or reservation of resources, such as compute resources on a node (e.g., on a CPU or on a dedicated compute block such as an FPGA or GPU), specific device capabilities (access to sensors / actuators, security devices (e.g., TPM), pre-installed software), storage resources on a node (memory or disk), network resources (latency or bandwidth potentially guaranteed via TSN), and so on.

[0180] Extended orchestrator rule sets can be defined to include criteria beyond standard compute, storage, and memory metrics, such as to specify application cycle time, application runtime, application input / output signal dependencies, or application process order (e.g., a mandatory order that specifies which application(s) run before or after other application blocks). This orchestration technology can provide, at the distributed application control policy level, the ability to leverage low-cost commodity hardware and software to achieve better system performance at the control policy level while simultaneously achieving new levels of system redundancy and failover across multiple applications running at ISA levels L1-L3 at a lower cost. Furthermore, orchestration sensitivity at a broader control policy level can enable new levels of high availability for embedded systems at a lower cost. This can result in increased uptime for general systems and applications for orchestrated and coordinated control applications, while reducing unplanned downtime for production operations at higher ISA levels compared to traditional approaches.

[0181] For systems that have system redundancy designed into their automated configurations, the following orchestration technologies can also enable additional maintenance tasks to occur (without production downtime). These technologies enable improved interoperability when executing control policies across vendor hardware that leverages platform-agnostic virtualization and containerization. These technologies also leverage current, historical, and simulation results to optimize workload placement for operational technology environments used in real-time operations. Furthermore, these technologies can leverage predictions of future orchestration events to pre-plan workload placement.

[0182] In the example, a distributed resource pool is defined as a combination of compute, storage, and memory across networked computing assets, with added functional block scheduling frequency and latency tolerance (before and after processing dispatch) for the purpose of executing application control policies. For example, a control policy (or application) can be defined by a physically distributed, coordinated set of building blocks with very strict timing, block-to-block scheduling, and runtime requirements for execution. The temporal arrangement of these building blocks coordinates the execution order, processing latency, and complete execution cycle of all building blocks that make up the entire application control policy.

[0183] Figure 10 The diagram illustrates an orchestration arrangement of an example cascade control application 1040 based on a configuration of distributed system building blocks 1010. Specifically, the diagram depicts an example building block set 1005 based on the IEC 61499 function block standard. Figure 10 The application shown demonstrates a general layering strategy employed in modern distributed control systems. For this example, a subset of all application blocks (block 1010) is illustrated for illustrative purposes; however, all application blocks shown may be included as dependencies for a specific implementation.

[0184] for Figure 10 In the example control application 1040 shown in FIG, function blocks A, B, C, and D (1022, 1024, 1026, 1028) are arranged in a cascade control design within a control subsystem. Each general-purpose building block (independent function block or application) executes a designated algorithm as part of a distributed control strategy to control an output (flow valve 1030). In this example, the control function block output is sent to the next function block as its input value. When a particular block is taken offline or "shed" due to some system anomaly, the link to the subordinate building block is handed over to the operator for manual control.

[0185] For the cascading strategy to work, the application cycle time, application runtime, application input / output signal dependencies, and application process order must be maintained for each block of the control loop. When these links are lost in production, significantly reduced efficiency ensue, representing a major inherent loss at the industrial level. The definition of an extended orchestrator rule set using this technique can address each of these resource issues.

[0186] The layering of capabilities within the expanded orchestrator rule set enables the addition of more advanced algorithms that directly impact production costs, improve product quality, and enhance process efficiency, while protecting worker safety through a loosely coupled set of design principles that enables individual applications to be taken offline and downgraded to lower levels of control to protect the overall operation. Without this layering of application control, new solutions would be difficult to implement and operations would be more prone to incidents. Furthermore, orchestrating these application assets at the control policy level further improves overall uptime and system performance, which directly benefits manufacturing and process operations.

[0187] Conventional IT orchestration strategies typically provide the ability to dynamically move individual application assets (function blocks) within the system; however, in this example, the coordination of distributed function block applications is orchestrated across all function blocks that define a specific control strategy. Collective function block links and associated state information are maintained to orchestrate these building blocks across system resources, allowing applications to remain online and avoid falling out of more fundamental security controls.

[0188] Figure 11 An example application distribution map depicting a control policy for an orchestration scenario including four applications, where application redundancy is depicted for native, virtual machine, container, and container within virtual machine deployments in design 1120. As illustrated, orchestration of application assets may include different deployment options to consider for dynamically allocating resources, subject to various compute, storage, memory, and application constraints.

[0189] Note that for Figure 11In the illustrated scenario, the applications (Applications 1 through 4) defined in orchestration scenario 1110 are specified to run at different frequencies. In this example, loops and runtime dependencies are the primary factors in the runtime orchestration decisions. Specifically, in the depicted example, Application 1 can be orchestrated within a 30-minute window and control policy execution is preserved; Application 2 can be orchestrated within a 5-second window and control policy execution is preserved; and Applications 3 and 4 can be orchestrated within a 1-second window and control policy execution is preserved. If the execution window for orchestration is missed, the application link is broken and the control policy is downgraded to a SAFE state until operations close the loop again.

[0190] Figure 12 The diagram illustrates example orchestration scenarios 1210A and 1210B suitable for handling function block application timing dependencies. As shown, in addition to more standard resource metrics that define where applications can be deployed to maintain error-free operation, application cycle, runtime dependencies, and current state also play an important role. For example, a control strategy that executes with relatively slow cycle times and frequencies can be run on a device with lower computing resources and does not need to be co-located with other dependent application blocks of the control strategy. Conversely, it may be necessary to co-locate all applications that need to execute with very fast cycle times and frequencies on the same device to ensure error-free operation of the control strategy.

[0191] exist Figure 12 In the example, orchestration scenario 1210A illustrates a scenario where applications 1-4 (application deployment 1230A) can be distributed across multiple independent nodes of a system to execute process 1220A. In contrast, orchestration scenario 1210B illustrates a scenario where applications 1-4 (application deployment 1230B) cannot be distributed across multiple independent nodes of a system due to cycle and runtime constraints. Instead, applications 1-4 must be orchestrated together for any orchestration event to successfully execute process 1220B.

[0192] Figure 13 Depicted are example orchestration asset deployments showing various deployments of orchestration assets (applications 1320) under the control of orchestrator 1310. Specifically, the example diagram illustrates a potential dynamic application outcome based on available system resources. As depicted, the examples cover VM, Container, VM+Container, and native node deployments. Figure 13 In the example, nodes 1, 6, 10, and 14 are active, showing how different applications in the same orchestration can operate in different system deployment types.

[0193] Figure 14A flowchart 1400 depicts an example orchestration sequence for a distributed control application strategy. In this example, each function block application resides in a different computing node of the system. Specifically, Figure 14 An orchestration method is implemented that considers application cycle time, application runtime, application input / output signal dependencies, and application process order for each block of a control loop, in addition to compute, storage and memory, and network resource availability, to effectively allow orchestration of control applications across available resources without interrupting control execution.

[0194] The arrangement of the various building blocks (or functional blocks) occurs in the Figure 14 The depicted complete control strategy is applied within the boundaries of the defined boundary conditions. Furthermore, current state and historical information, when combined with a defined set of constraints applied to individual functional blocks, provides a means for performing various multi-level optimization methods for resource allocation, which can also include predictions of orchestration possibilities for a wider range of control strategies. By combining prediction and constraint management with real-time optimization, a high level of resilience can be achieved for embedded infrastructure.

[0195] In an example, the operation for monitoring the functional blocks of the distributed control application (operation 1410) may include monitoring various forms of current and historical state data. This may include monitoring: available computational overhead; available computational speed; available storage; available memory; application cycle time; application runtime; application link dependencies; application process sequence dependencies; or application-specific orchestration errors.

[0196] In yet another example, the operation for updating the prediction (operation 1420) may include: orchestration optimization = f(current state data, historical state data, constraints) for each control policy; orchestration optimization = f(current state data, historical state data, constraints) for each application building block; or, orchestration prediction = f(current state data, historical state data, constraints) for each application building block.

[0197] In a further example, operations for detecting system anomalies for application building blocks may be evaluated (operation 1430). These may be subject to constraints defined for each application, such as: computational overhead allowance; computational speed minimum requirement; storage minimum requirement; memory minimum requirement; application cycle time limit; application runtime limit; application link dependency for input and output dependencies; application process sequence dependency for input and output variables; and application system error triggering for orchestration events.

[0198] In another example, the operation can evaluate whether orchestration of any functional blocks is required (operation 1440). For example, if application constraints l..n are violated, then an orchestration event is required. In another example, the operation can also evaluate whether control strategy orchestration is feasible (operation 1450). This can evaluate whether an application needs to be moved to another node within defined constraints, whether multiple applications need to be moved due to application dependencies, and if so, whether and how the group of applications can be distributed. In another example, if orchestration is not feasible, a degradation or disconnection control strategy can be implemented (operation 1460), and the active functional block profile can be updated accordingly (operation 1480).

[0199] In yet another example, in response to verifying that orchestration of the functional blocks is required and that the control strategy orchestration is feasible, an operation is performed to orchestrate the building block application of the control strategy (operation 1470). If the orchestration is successful, this results in a predicted reset (operation 1490). If the orchestration fails, this results in using a degraded or disconnected control strategy (operation 1460) and updating the active functional block profile (operation 1480).

[0200] Figure 15 The figure shows a flow chart of an example method for orchestrating distributed mission-critical workloads and applications using distributed resource pools. Based on the previous example, this method can be implemented based on (for example, as referenced Figure 10-13 The method also provides the ability to dynamically orchestrate distributed applications and groups of dependent applications using an extended set of application-specific dependencies (described above). The method also enables the ability to dynamically analyze and simulate network bandwidth before committing to an orchestration policy. The method also provides the ability to predict orchestration events before they occur and proactively plan potential optimized resource placement for control policy workload orchestration.

[0201] In flowchart 1500, example operations include identifying application-specific dependencies (operation 1510); dynamically creating orchestration groups of distributed and dependent applications based on the identified dependencies (operation 1520); and predicting orchestration events using the orchestration groups (operation 1540). In an example, predicting orchestration events includes dynamically analyzing and simulating network bandwidth (or other resources) in an example scenario (operation 1530), and analyzing the occurrence of orchestration events in the example scenario.

[0202] Based on the predicted orchestration events, operations can be performed to define and modify the extended orchestrator logic rule set. These operations can also include detecting the predicted orchestration events (operation 1550), and optimizing resource placement based on the predicted orchestration events (operation 1560). For example, referring to Figure 14 The techniques discussed can be incorporated into various aspects of a changed orchestration strategy.

[0203] Choreography for traditional (brownfield) environments

[0204] Orchestration is the act of matching user requirements for an application (which may consist of many processing, networking, and / or storage components) with the capabilities of the physical world and deploying the application (such as by configuring the physical world and allocating and configuring the application components). Orchestration is often applied in enterprise environments to deploy highly scalable services into the homogeneous and virtualized environments in which these applications are designed to operate.

[0205] When orchestration is applied to IoT environments, especially those with existing ("legacy") devices (e.g., "brownfield" deployments), the problem changes in several ways: there are typically a large number of devices; the set of target devices is highly heterogeneous; some application components may not be designed for orchestration; some hardware devices may not be designed for orchestration solutions; and some devices may be proprietary and closed / fixed-function devices.

[0206] Using conventional approaches, software modules must conform to specific APIs to be orchestrated so that they can be correctly deployed and configured. Typically, to receive the orchestrated software, hardware nodes run orchestration software and provide specific APIs to the executing software. Consequently, several issues can arise, including: how to scale across many devices and applications; how to allow orchestration of software modules that were not designed to be orchestrated in an unmodified state; how to allow orchestration of software modules onto legacy hardware nodes or hardware nodes that are not originally capable of supporting the orchestration stack; and how to self-monitor a heterogeneous set of physical nodes to manage resource utilization.

[0207] In an example, the following techniques enable orchestration of choreography-unaware code by encapsulating the code within an orchestration-aware shim. In a further example, the following techniques enable choreography-unaware devices to participate in orchestration. In contrast, existing approaches do not consider the choreography problem (particularly end-to-end choreography in heterogeneous environments), which introduces significantly different issues and requirements. Similarly, in a further example, choreography self-monitoring can be used to achieve self-reliant and self-organizing choreography that learns from failures and incorporates feedback into better choreography methods.

[0208] Orchestration techniques allow the dynamic deployment of individual software components of distributed IoT applications across a collection of available hardware resources, taking into account resource capabilities and application requirements and constraints. Current orchestration techniques tend to assume that (1) software components are designed to be orchestrated by implementing an orchestration API, and (2) devices are designed to receive orchestrated software by providing orchestration middleware. The techniques herein enable the orchestration of legacy software components by wrapping such legacy components inside orchestrated components, which provides a plug-in architecture for interacting with legacy software using standard or custom mechanisms. For example, a standard plug-in can receive a communication port number from an orchestration API and set the port number on standard software such as a web server through a configuration file or environment variable. Similarly, custom plug-ins can be written to support proprietary software, for example.

[0209] Figure 16A The figure illustrates an example scenario of the orchestration between an orchestration engine 1610A and associated modules. As shown, the orchestration engine 1610A deploys two orchestration modules 1620A and 1630A. The two modules each use an orchestration API (respectively 1640A and 1640B) to receive configuration parameters from the orchestration engine 1610A. For example, if module 1 1620A is an http client and module 2 1630 is an http server, module 1 1620A can receive endpoint information, such as an IP address and a port number, that the module needs to communicate with module 2 1630. In some cases, the orchestration engine 1610A provides the port number that module 2 1630 should bind to to module 2 1630, while in other cases, module 2 1630 can provide communication information to the orchestration engine after binding. In either case, the two modules all use APIs (e.g., APIs 1640A and 1640B) to establish communication parameters and become connected.

[0210] Figure 16BThe figure illustrates an example scenario of orchestration between an orchestration engine and associated modules (including legacy modules). Orchestration engine 1610B deploys two different modules (1620B and 1660), one aware of orchestration and the other a legacy module that is unaware of orchestration. In this case, orchestration engine 1610B also deploys a shim 1650 along with legacy module 1660. This shim 1650 understands any custom configuration mechanisms associated with legacy module 1660. For example, if legacy module 1660 is an Apache web server, shim 1650 can be configured to negotiate a port number through orchestration API 1640D, and then configure the web server's port number using a configuration file, command line parameters, or environment variables (or similar mechanisms) before starting the Apache server. The client in schedulable module 1640C will use schedulable API 1640 to negotiate client communication parameters to operate in the same manner as the previous example, and will therefore be able to connect to the Apache web server.

[0211] In an example, workloads performed by traditional hardware devices can be processed in a similar manner by pairing each traditional device with an orchestration device. As an example, Figure 17A The diagram illustrates an orchestration scenario with an orchestratable device. An agent on an orchestratable device 1710A collects information about the device's available resources and reports it as telemetry to an orchestration engine 1720A. A typical orchestratable device can indicate its capabilities to the orchestrator, receive workload execution requests from the orchestrator, and execute workloads 1730A. Traditional devices do not support these capabilities and are therefore paired with an orchestratable device.

[0212] As a further example, Figure 17B The figure shows a scenario of orchestration with traditional devices 1780. Each orchestratable device (e.g., device 1750B) represents the capabilities of the traditional system to the orchestration system. When the orchestration system requests a workload on the traditional system, the paired device is responsible for executing the function on the traditional device. This can take the form of a remote procedure call or a custom API. As a result, the orchestration engine 1720B is able to match and deploy the appropriate workload to the device. For traditional devices, the agent on the orchestratable device 1750B paired with the traditional device 1780 can discover the existence of the traditional device 1780 and measure the capabilities of the traditional device (e.g., via an RPC mechanism). This information is then passed by the agent as telemetry 1740B to the orchestration engine 1710B. When the orchestration engine 1720B passes the workload 1730B for the traditional device 1780, the agent 1760B deploys it to the traditional device 1780 (e.g., via an RPC mechanism). Therefore, the encapsulation mechanism allows both traditional hardware and software to participate in modern IoT orchestration solutions.

[0213] Orchestration technologies typically provide for the scheduling and management of a flat set of resources. Orchestrated resources can include compute (physical or virtual devices), networking (physical or virtual interfaces, links, or switches), or storage (databases or storage devices). Orchestration can take the form of task (execution unit) orchestration, container orchestration, virtual machine orchestration, network orchestration, or storage orchestration. Or it can be all of these simultaneously, in the form of end-to-end application orchestration.

[0214] Figure 18 A coordinated scenario of workload orchestration in a single-level orchestration environment is depicted. The single-level orchestration environment shows a scenario where all platforms participate equally in orchestration: each node (e.g., nodes 1830A, 1830B, 1830C) describes its available resources to an orchestration engine 1820 (typically centralized, such as at orchestrator 1810), which performs the scheduling function by sending telemetry, and orchestration engine 1820 dispatches subsets of multiple nodes to run multiple parts of the overall application workload. Thus, as Figure 18 As shown, various workloads (1821A, 1821B, 1822, 1823A, 1823B) are assigned to various nodes 1830A, 1830B, 1830C and executed using corresponding agents 1840A, 1840B, 1840C. This approach provides a flat orchestration structure and implies a minimum level of capability for each node 1830A, 1830B, 1830C, so that each node can fully participate in the orchestration process.

[0215] However, an orchestration can be made hierarchical by breaking it down into various functions and functional operations. Figure 19 An example functional hierarchy for orchestration is depicted, illustrating how application orchestration 1910 provides the top-level domain of control for orchestration. If end-to-end application orchestration is done at the top level, the details of network orchestration 1920, virtual machine orchestration 1930, task orchestration 1940, and storage orchestration 1950 can be delegated to sub-orchestration modules. The sub-orchestrator can be used to determine how to optimize each sub-problem and configure resources in each sub-domain.

[0216] Figure 20 The diagram illustrates an example deployment of a general-purpose layered orchestration solution. Figure 20 The deployment in FIG. 2 depicts a general hierarchical structure of sub-orchestrators 2040A, 2040B, and 2040C, where a pool of orchestratable devices can be called upon to implement multiple parts of the overall application.

[0217] In the example, each child orchestrator (e.g., 2040A-2040C) receives telemetry from the orchestratable devices (e.g., 2050A-2050G) in a given pool of orchestratable devices. The telemetry indicates the resources available in the pool. The child orchestrators aggregate the telemetry and forward it to the top-level orchestrator 2010. The top-level orchestrator receives telemetry from the child orchestrators (2040A-2040C), which inform the top-level orchestrator 2010 of the total resources available in the pool. The top-level orchestrator 2010 then dispatches a subset of the total workload to the orchestration engine 2020 based on the telemetry. The child orchestrators, in turn, schedule the subset of the workload onto each orchestratable device in the pool. Note that while two levels of orchestration are used in this example, additional levels can be implemented.

[0218] In some scenarios, it is possible for the orchestrator 2010 to oversubscribe resources in a pool, assuming that resources can be shared across time and space (between pools). In addition, the child orchestrators may be able to borrow equipment from underutilized pools to temporarily handle excess load. For example, in Figure 20 In the example shown, if cluster 1 2030A becomes overloaded, one or more slave devices can be borrowed from cluster 2 2030B or cluster 3 2030C.

[0219] Although Figure 20 The approach described in assumes that all devices are programmable, but in reality, many programmable devices can be very low-cost microcontrollers with minimal memory and storage. Each group of potentially hundreds or thousands of these low-cost sensing solutions can be controlled by a more powerful device. To address this situation, Figure 21 The figure shows an example of using a hierarchical orchestration provided by a slave node.

[0220] Figure 21 The scenario provides a similar approach to that discussed above in the hierarchical orchestration scenario, where the master orchestration device 2110 can represent the capabilities of many other slave nodes. Such capabilities may include, for example, the ability to sense from a specific sensor device, or the ability to perform calculations using a specific FPGA component. The agent reports these capabilities to the orchestrator 2110, which dispatches the workload to the individual master orchestration device. However, the master node does not necessarily run the workload, but can instead send it to a slave node (e.g., nodes 2150A-2150H) that has the individual capabilities required for the workload. This process occurs transparently to the orchestrator 2110, which only cares about the work being executed.

[0221] To implement this master / slave relationship, some simple primitives are implemented on the slave nodes, including: (a) detection, such that the presence of the slave node must be detected by the master node, and the failure of the slave node (and therefore the deployed workload) must also be detected; (b) discovery, such that the available resources on the slave node must be discoverable by the master node, and such information helps determine the type and number of workloads that can be deployed; (c) deployment, such that the master node is able to deploy workloads on the slave node (e.g., RPC, firmware deployment, etc.).

[0222] Figure 22 The diagram illustrates an example workflow for a slave node used in a hierarchical orchestration scenario. In the example, the node will wait to receive a discovery request from the master orchestration device that leads its cluster (operation 2210). During this time, the node may be waiting in a lower power state. The request may include some kind of cryptographic challenge, such as a nonce that will be encrypted by the slave node. When the slave node receives the request, the slave node can send back some credentials (operation 2220) to prove that the node belongs to the cluster. For example, the node can encrypt the nonce with a private key and send back the result. The slave node can also send telemetry to the cluster leader in the form of a capability set (operation 2230). The slave node will then wait for its instructions (operation 2240), which may be in the form of a workload to be executed. When the slave node receives the workload (operation 2250), the slave node may need to reprogram itself (operation 2260), perhaps by refreshing its programmable memory; after reprogramming, the slave node can then continue to execute the workload (operation 2270).

[0223] Scheduling in a layered solution can introduce complexities. For example, an agent on a master node must be careful to correctly describe the relationships between the resources it represents so that the orchestrator does not mistakenly believe that resources are co-located on the same node when they are actually distributed across many slave nodes.

[0224] The layered orchestration mechanism described above allows for the creation of dynamic SDIS solutions that are more heterogeneous in nature, including the use of resource-constrained (and inexpensive) components that would otherwise not be able to fully participate in orchestration. Furthermore, this deployment allows for a smaller number of IA-based nodes (expensive resources) to be used as master nodes, providing orchestration mechanisms for each cluster.

[0225] In a very large IoT framework, while the solution chosen by the orchestrator regarding what software to deploy on specific hardware components may initially be correct, this may change over time. Furthermore, the overall system capacity must be monitored to ensure that available resources are not exhausted. Therefore, the overall solution needs to be monitored for software and hardware issues, such as CPU overload, and appropriate steps taken to address them. The following techniques enable self-reliant and self-organizing orchestration that learns from failures and incorporates feedback into better orchestration.

[0226] In a further example, a control loop-like checking and feedback mechanism can be added to the orchestration approach discussed above. Individual components, including software, networking, storage, and processing, may have built-in monitoring mechanisms or may require frequent polling to enable such management. This can be provided by extending the orchestration layer, which tracks all available resources, to include tags indicating what operations need to be monitored, such as CPU, memory, application response latency, application behavior, network latency, network bandwidth, and specific hardware.

[0227] Figure 23 The diagram illustrates an example configuration of a monitoring and feedback controller 2310 suitable for coordinating and implementing orchestration self-monitoring functionality. In the example, monitoring and feedback controller 2310 collects software data 2320, hardware data 2330, and network data 2340 from various client nodes 2350, 2360. These client nodes 2350, 2360, in turn, operate orchestrated operations and workloads under the guidance of orchestration server 2370.

[0228] In this example, client nodes 2350 and 2360 are monitored for hardware and software overloads. For example, if a device's CPU or memory reaches 50% capacity, the device can be closely monitored. If capacity reaches 80%, the device can be replaced, or the workload can be migrated to a device that better matches the workload being executed. If hardware dependencies exist, additional nodes can be added to handle the software load. In a similar example, network traffic can also be monitored. If a large amount of unknown traffic is observed, or if less than expected traffic is observed, the system can check the performance of the client node. Such checks can also indicate hacker intrusion or loss of network connectivity.

[0229] The monitoring and feedback controller 2310 enables looping back to the management server to dynamically control behavior. Feedback loops are used not only for client nodes but also for servers. For example, a monitoring mechanism can monitor the performance of server nodes, perhaps by monitoring network traffic between servers. For example, if the gossip protocol that monitors server cluster health and ensures that the elected leader is always available is consuming too much bandwidth, the protocol parameters can be dynamically modified to better adapt to the current number of servers, link conditions, and traffic levels.

[0230] The monitoring and feedback controller 2310 may also incorporate logic for learning from failures. If a node is consistently failing, there may be an underlying hardware issue, and the node may be scheduled for maintenance. If many nodes are failing in a particular physical area, this pattern may be detected, and those nodes may be scheduled for manual inspection. In a further example, nodes may monitor themselves for potential issues. For example, each node may monitor itself, rather than requiring a monitoring solution to poll individual nodes to determine their health. For example, if a node is running low on memory, it may report its condition to a central monitor.

[0231] Other aspects of self-monitoring and management can also be incorporated along with orchestration. The system can specify maintenance schedules for individual nodes. Each node can be scheduled for service after a certain number of operating hours, at which point the node will be removed from the set of nodes available for scheduling by the scheduler.

[0232] In a further example, the self-monitoring functionality can also provide capacity planning. For example, if networking or processing usage approaches capacity, operators can be notified to increase capacity. The system can help operators plan by specifying how many resources are needed and what kind of resources are needed. For example, the system can specify that additional nodes are needed on which tasks can be deployed, and that these nodes should have a certain minimum memory and storage capacity. This self-monitoring functionality allows orchestration solutions to be highly scalable and easily adapted to the infrastructure.

[0233] Figure 24 A flowchart 2400 of an example method for orchestrating devices in a traditional setting is illustrated. As shown, flowchart 2400 includes a series of end-to-end actions for configuring and operating an orchestration in a brownfield environment, which features establishing communications with traditional components, establishing an organized orchestration, and operating, monitoring, and adjusting the orchestration. It will be understood that flowchart 2400 is provided at a high level for illustrative purposes, and that additional configuration and usage operations described above can be integrated into the operational flow.

[0234] As shown, flowchart 2400 includes operations for establishing an orchestration shim to configure a legacy software module (operation 2410), transmitting the configuration to the legacy software module via an orchestration shim API (operation 2420), and collecting telemetry from the legacy hardware device via an orchestratable device agent (operation 2430). Further configuration operations (including Figures 16A-17B The operations depicted and discussed in the ) may include the configuration of programmable hardware devices and programmable software modules.

[0235] As also shown, flowchart 2400 includes operations for organizing a hierarchical structure of components (operation 2440), such as configured traditional components and programmable components. The organization may include organizing the components into various hierarchical structures (operation 2450), performing detection, discovery, and deployment of various slave node components (operation 2460). Further detection and hierarchical organization may also occur (including in Figures 18-22 operations depicted and discussed in

[15] .

[0236] As also shown, flowchart 2400 ends with operations for distributing workload to various components in the hierarchy of components based on telemetry and other configuration data from the components in the hierarchy (operation 2470) (including Figures 18-22 2400 is further used to communicate between components of an organized (layered) orchestration (including Figure 23 ) to collect and monitor software data, hardware data, and network data to allow self-monitoring and configuration changes (operation 2480); in response, an orchestrator, administrator, or other entity may provide feedback and control to various components of the organized orchestration (operation 2490).

[0237] Self-describing orchestration components

[0238] In the development of industrial solutions, engineers can design the solution as a diagram of modules that can be deployed into an IoT system. Figure 25 The diagram illustrates an example industrial control application scenario, specifically depicting the problem of maintaining the temperature of a water tank 2530 by heating the surrounding oil jacket with a heater 2536. The water and oil temperatures are monitored by respective sensors 2532, 2534 to control the process. A set of computing nodes 2520, some of which can be connected to physical sensors and actuators in the system, may be available on which software modules can be deployed.

[0239] In this example, a control engineer can design a control system application 2510 to perform a functional operation, such as controlling temperature as a cascade control loop consisting of a graph of software modules that can be deployed on available computing nodes. A sensor module can read data from a primary sensor 2532 that reads a value from a sensor in the water. This value is fed into the input of a PID (proportional-integral-derivative) controller module (e.g., a controller having one or more proportional, integral, or differential control elements) that attempts to meet a specific set point. The output of the PID controller is fed into a scaling module, the output of which establishes a set point for another PID controller. The second PID controller receives its input from a module that reads data from a sensor in the oil (e.g., from sensor 2534). The output of the second PID controller is sent to an actuator module that controls a heater element 2536. In the example, any one of the PID controllers can be a controller that incorporates proportional, integral, or differential control (alone or in any combination) as part of any number of functional operations.

[0240] To properly deploy such a configuration, a control engineer describes the control application and the functions and operations within the control application. The following discusses techniques for defining a configuration for a language used to describe control system applications. The following further discusses the use of self-describing modules upon which control system applications can be implemented, and an orchestrator that can leverage the language and self-describing modules to deploy efficient solutions to compute nodes.

[0241] The following approach specifically implements the use of self-configuration and self-describing modules for enhanced implementation of orchestration in the SDIS environment discussed herein. As discussed herein, self-describing modules allow for a better understanding of which platform resources need to be deployed and which platform resources make orchestration easier by clarifying requirements or constraints. Self-describing modules provide for the separation of the self-description of a module from the self-description of an end-to-end application. Self-describing modules also provide the ability to express multiple alternative implementations of a given software module and the ability to balance between multiple implementations. This approach can be implemented in an architecture for automatically evaluating the compromise between alternative implementations of a module and an application, thereby helping users orchestrate optimized applications on IA (instruction architecture, such as x86, ARM) devices.

[0242] In the following examples, a module is a component of an application that is deployed by an orchestrator. A module has a module manifest (e.g., Figure 13 , and referenced in the examples in Table 1). An application consists of a collection of modules whose inputs and outputs are connected together. Using an application specification (e.g. Figure 26As shown in Figure 2 and referenced in the example of Table 2, an application is described. In this example, the application specification is created by the user to define the end-to-end application. The application specification provides input to the orchestrator, along with any applicable module manifests. The application specification can also be used to specify the modules, their interconnections, and any additional requirements that must be met when deploying these modules. Thus, using the module manifests and application specifications in this manner enables the functional operation of the end-to-end application to be realized and implemented.

[0243] Attempts have been made in many settings to define the concept of an end-to-end application for application deployment; however, existing orchestration approaches focus on IT considerations and do not provide a flexible approach for use in industrial systems. This approach does not consider end-to-end applications, encompassing everything from edge devices to cloud deployments. Furthermore, existing orchestration systems do not allow users to express alternative implementations for a given software module or provide users with a means to evaluate or express trade-offs between multiple alternative implementations. The following self-describing modules and self-describing language enable a better understanding of which platform resources need to be deployed and, therefore, make orchestration easier and more accurate by clarifying the appropriate requirements or constraints.

[0244] In an example, in addition to the self-describing modules on which control system applications can be implemented, the SDIS implementation can be extended to provide a language for describing control system applications. From these two elements, the orchestrator can deploy working solutions to various computing nodes and resources. Thus, the techniques described herein provide: (1) a mechanism for establishing self-descriptions for orchestratable modules in order to separate the end-to-end application from the individual modules; (2) a mechanism for allowing the system to dynamically select between multiple alternative implementations of a module for deployment; and (3) a mechanism for allowing the system to reason about which alternative implementation is best in different situations.

[0245] Figure 26Depicted is an overview of a control application represented by an example control application diagram 2600, which is represented at the sensor and actuator level. As shown, a control application is defined by a control engineer as a diagram of software modules, where the output of each module (e.g., the output from sensor A 2610 and sensor B 2620) is connected to the inputs of other modules (e.g., the inputs to actuator C 2640 and PID controller 2630). The control engineer may also specify other factors, such as starting values ​​for module parameters. The control engineer may find these software modules in a software library or request that the IT department implement a custom module. In an example, the diagram may be defined using a graphical user interface or other visually based representation. For example, the example control application diagram 2600 may be defined by a control engineer to reflect the inputs, outputs, and controllers of an industrial system. The example control application diagram 2600 may reflect the connections of the physical system and be used to perform the various functional operations (as well as actual changes, measurements, and effects) of the control application.

[0246] Figure 27 Describes a method for implementing self-describing control applications such as Figure 26 An example software module definition for a control system module (PID controller 2710) depicted in FIG. In the example, the code for this software module is written with several assumptions, including that the module is unaware of the node on which it will be deployed and that the module can communicate with adjacent modules through a set of named interfaces. Interfaces can be directional to allow for connection-oriented protocols (which typically have client and server endpoints), which are typically established in a directional manner, but not necessarily in terms of the direction of data flow (which can flow in one or both directions).

[0247] In a further example, the module's code has requirements (e.g., network requirements 2740) for the channel over which it communicates with adjacent modules (bandwidth, latency, jitter, etc.). However, the module does not know what modules it will communicate with or what nodes these modules will be deployed to. The module does not know the communication parameters for its communication endpoints or other communication endpoints. The module may require a certain amount / type of processing resources, memory resources, and storage resources, and may require other hardware and software dependencies (libraries, instruction sets, chipsets, security coprocessors, FPGAs, etc.). In addition, the module may allow a set of named starting parameters to be specified (e.g., parameters 2720).

[0248] To make the code self-describing, module developers can create a module manifest for use with the software module, where the module manifest is used to identify and describe key characteristics of the control environment used to execute the software module. In an example, the characteristics can include features such as: (a) communication interfaces (of the PID controller 2710), including the name, type (client, server, publish / subscribe), protocol (dds, opc-ua, http), or QoS requirements (if any) for each interface; (b) parameters and default starting values ​​(e.g., control parameters 2720); (c) platform requirements (e.g., instruction set, OS, RAM, storage, processing) (e.g., requirements 2750); (d) dependencies (e.g., libraries, hardware, input signals, etc.) (e.g., dependencies 2730); (e) deployment requirements (security, isolation, privacy, orchestration style); (f) signature of the code module (e.g., signature 2760).

[0249] List of example modules for control system applications and Figure 27 The modules executed in can be represented by the following definitions:

[0250]

[0251]

[0252] Table 1

[0253] In a further example, a control engineer may utilize a library of one or more software modules to create or define a control system application. For example, a graphical user interface (GUI) may be used to design a diagram of a control system application (e.g., similar to Figure 26 The GUI can utilize a module list to indicate the details of each code module and how the various code modules can be connected to each other. In addition, the user can utilize drag-and-drop and other graphical indication methods to select appropriate modules and connect and configure them to design and Figure 26 The control application diagram shown is similar to the diagram.

[0254] The result of this information compiled into an application specification for a control system application may be encoded in an application specification format similar to the following example:

[0255]

[0256]

[0257]

[0258] Table 2

[0259] An application specification defined in this way allows the control engineer to: select a set of modules to use, specify values ​​for parameters beyond any default values, specify any additional constraints or resources beyond those specified by the modules themselves, and specify how the modules are to be linked together. Additionally, the application specification can assign specific parameters to the link, such as assigning a topic name to a publish / subscribe channel or a port number to a server endpoint (making the communication endpoint accessible from outside the application).

[0260] In an example, an application specification may also specify alternative implementations for the same functionality in the application (e.g., each version of the functionality implemented by a different module). For example, consider two versions of a module that implement the same functionality for two different hardware architectures. Module writers can specify these alternatives in the module manifest, such as shown in the following example:

[0261]

[0262] Table 3

[0263] In another example, a controls engineer might specify these overrides in the application specification as follows:

[0264]

[0265] Table 4

[0266] In this example, the orchestrator can be implemented by selecting an appropriate software module and deployed on a node of one of the two architectures (x86 or ARM) that satisfies either of the two constraints.

[0267] The use of self-describing module representations can be applied to other classes or types of resources. For example, such self-describing representations can be applied where an algorithm can be implemented on a general-purpose CPU, GPU, or FPGA. In this case, a score can also be provided in the application or module specification to indicate which module is preferred. The score can be both algorithm-specific and data / application-specific, and therefore requires some knowledge on behalf of the developer or control engineer. In addition, the use of scores can enable the control engineer to optimize the selected control application by utilizing software modules (if available) that are optimized for a specific LA hardware platform, such as an FPGA or neural network processor (NNP).

[0268] The use of self-describing module representations can be further generalized to consider more general resources. For example, a first version of an algorithm can be optimized for memory resources, while a second version can be optimized for storage resources. In this case, the first version has small memory resource requirements and larger storage requirements, while the second version has large memory resource requirements and small storage requirements. The orchestrator can select modules based on the resources available on the available set of nodes. Furthermore, scoring can help determine which module is preferred without constraining other factors.

[0269] Self-describing representations can also be used in the case of node associations, for example, in the case where module A is deployed on node A with a preference level of N, while module B is deployed on node B with a preference level of M. If N indicates a higher preference than M, the system will attempt to deploy module A to node A (if available), otherwise deploy module B to node B.

[0270] However, one of the challenges of self-describing characterization is that the control engineer may not actually know which version of a given software module performs a certain application function most efficiently, or even what criteria can be used for that software module to produce the best end-to-end results. The control engineer may only observe objective results (e.g., which solution "seems most responsive."). With many combinations of software modules, criteria, and options, the framework can be used to test which combinations of system modules and alternative implementations are effective.

[0271] Figure 28 Describes an architecture for automatically evaluating alternative implementations of software modules. Specifically, Figure 28 The architecture provides a framework for simulating various combinations of modules from an application specification and characterizing the results. Various data from the user's application specification and module manifest 2820 are provided to the system. The system has access to all module images stored in the module image repository 2810. Each module can have several alternative implementations.

[0272] In the example, a series of experiments are performed and evaluated on various combinations of these implementations. The experiments can be controlled by a characterization controller 2830, which will ensure that the various combinations are executed. The experiments will work in conjunction with an orchestrator 2840, which is responsible for deploying the modules specified in the application specification and module list 2820 onto a set of simulators 2850. The simulators 2850 simulate the hardware defined by the given alternatives specified in the application specification or module list 2820 (for example, a specific FPGA or CPU with a certain amount of available memory). The orchestrator 2840 will deploy the application, interconnect the components, and run the application. The system will then automatically score the system based on a certain criterion (for example, end-to-end latency) using a score 2860, or the user can score the application based on a subjective criterion ("feeling agile"). Finally, the system will reason about the various combinations and determine the best combination to use, such as by utilizing a decision tree-based approach.

[0273] Figure 29 The diagram shows the Figure 28 Following the example depicted in FIG. 29, a flowchart 2900 of an example method for evaluating alternative implementations of a software module is provided. In flowchart 2900, an optional prerequisite includes an operation for determining a configuration of applications and modules that can operate within the system using application specifications and module manifest information (operation 2910). The prerequisite can be performed as a one-time event or repeatedly.

[0274] The operations of flowchart 2900 continue with defining and executing corresponding orchestration scenarios (operation 2920) by a characterization controller that is configured to execute an application module having one or more defined options in a simulator (e.g., a simulator configured according to a specific hardware setup) (operation 2930). Using the simulator, various modules and various module options may be executed, including using an alternative application module having one or more defined options in the simulator or another simulator configuration (operation 2940). Execution of the alternative application module may be repeated for multiple various software modules and multiple options.

[0275] The operation of flowchart 2900 continues by evaluating the results of the application module execution based on the defined performance metrics or standards (operation 2950). Then, the execution scenarios for one or more application modules are scored (operation 2960), ranked, or further evaluated using an automated or human-influenced scoring process. Based on the scores, various execution scenarios of the application modules can be incorporated or updated (operation 2970).

[0276] Figure 30AFlowchart 3000A of an example method for defining an application using self-describing programmable software modules is shown. The method begins with operations that define which software modules or application capabilities are selected and used as part of the application orchestration. These operations include the creation of a module manifest (operation 3010A) that describes the various characteristics of the orchestrated execution of modules of a control system application (e.g., an industrial control application in an SDIS). Further module definition operations also include defining various options and alternatives for the operation of various software modules (operation 3020A), and defining resource standards for the operation of various software modules (operation 3030A). These operations also include defining specifications for the application based on the definitions of the various software modules and the connection requirements and conditions for the features available within the various software modules (operation 3040A). Such definitions may include the above referenced Figures 26 to 28 Various operations discussed.

[0277] Flowchart 3000A continues with the simulation and evaluation of various software modules, such as described above with reference to Figure 29 The application setting of one or more simulations discussed above may be implemented in a different manner (operation 3050A). The output of the simulation may include priorities or other properties of various implementations of the modules. Based on the evaluation, specific combinations of software modules and options (priorities and other properties) for executing such software modules may be selected (operation 3060A), and these combinations may be deployed in an orchestrated application setting (operation 3070A). When combined with the constraints and properties of the physical system, such priorities and options may be used to inform the orchestration process.

[0278] Figure 30B A flowchart 3000B illustrates an example method for using self-describing, programmable software modules in an SDIS system implementation. In this example, the operations of flowchart 3000B are performed by an orchestration device (orchestrator) operatively coupled to multiple execution devices within a control system environment to execute the software modules. In this configuration, execution of a selected software module via at least one execution device implements the functional operation of one or more control devices within the control system environment. Furthermore, the orchestration device (orchestrator) can coordinate the execution of the selected software modules using an orchestration control policy within the control system environment.

[0279] Flowchart 3000B begins at 3010B with the optional prerequisite of creating a module list and an application specification listing the required system characteristics. Operation 3010B can be performed manually or through automated / computer-assisted features. The module list is used by the following processes to define the environment in which the software modules will execute the control system application.

[0280] Flowchart 3000B also continues at 3020B with an optional prerequisite of generating an application specification for a control system application, the application specification including matching module information and system characteristics (including parameters, values, etc., for execution). For example, the application specification for a control system application may define values ​​for control parameters of selected software modules, including indicating relevant connections or relationships between software modules or functions.

[0281] Flowchart 3000B continues at 3030B with identifying available software modules and, at 3040B, with identifying characteristics of a control system or control system environment from a module inventory. In this example, operational aspects of available software modules capable of performing specific functional operations within the control system environment are identified. The operational characteristics of the system identified in the module inventory may relate to one or more of: communication interfaces, startup parameters, platform requirements, dependencies, deployment requirements, or signatures.

[0282] Flowchart 3000B continues at 3050B with an operation of selecting one or more matching software modules based on the available software modules and system characteristics. For example, the selection may be based on a match between operational aspects of the available software modules and the identified operational characteristics of the system indicated in the module list.

[0283] Flowchart 3000B ends at 3060B with operations for executing the control system application, including the execution of associated software modules, according to the values ​​and characteristics of the application specification. Finally, flowchart 3000B includes operations at 3070B that allow for evaluation of the execution (or simulated execution) of the associated software modules, thereby allowing for further adjustments and feedback to the manifest or application specification. For example, the evaluation may include: evaluating the execution of the selected software modules in the control system environment using at least two different hardware architectures; and performing efficiency measurements on the operations executed using the at least two different hardware architectures. Other types of execution characteristics or deployments may also be evaluated.

[0284] In various examples, a control system application can be displayed and modified using a visual representation displayed in a graphical user interface. For example, the visual representation can be used to establish a relationship between one or more inputs or outputs and the control system application, including inputs or outputs involving the use of one or more sensors, actuators, or controllers.

[0285] Sensor bus redundancy

[0286] Sensor buses can have redundancy, such as using multi-layer field device redundancy in distributed control systems. Traditional industrial control systems use programmable logic controllers (PLCs) as key elements for controlling plant operations. A single PLC can communicate with and control hundreds of field devices and run control algorithms such as proportional-integral-derivative controllers. Due to the consolidated nature of PLCs, if a PLC fails, data from all downstream field devices will be unavailable, and the control functions being executed on the PLC will cease. A simple way to achieve complete resilience in an industrial control system is to deploy a fully redundant environment. However, achieving both is expensive and presents many logistical challenges.

[0287] In the systems and methods described herein, a field device abstraction bus (eg, Ethernet) is used, which decouples physical and functional requirements and improves scalability and can extend possible industrial architectures.

[0288] Addressing manufacturing process reliability and survivability. The Field Device Abstraction Bus enables any wired controller node in a distributed control environment to communicate with any wired field device. This "any-to-any" control architecture offers improved survivability by enabling healthy control nodes to assume the acquisition and control responsibilities of a failed control node. Healthy control nodes can be either control nodes with existing control responsibilities or "surplus" control nodes inserted into the system to improve survivability.

[0289] Expanded data availability. In existing systems, which are often proprietary and have tightly coupled functionality, data is often not freely available due to interoperability limitations. The implementation of a Field Device Abstraction Bus makes raw field data available to any authenticated consumer.

[0290] Previous solutions and architectures focused on consolidating capabilities into tightly coupled, single devices, exacerbating the "single point of failure" problem. Consequently, field device data, whether physical or virtual, was not located on the "bus." Only the host computer (PLC) could access real-time field device data, and if the host computer (PLC) failed, downstream field device data would be inaccessible.

[0291] The systems and methods described herein include a multi-layer field device redundant bus that enables an "any-to-any" relationship between controllers and field devices. The decoupling of controllers and IO enables simple failover and redundancy.

[0292] Improved system reliability and survivability is achieved by enabling any controller to access any field device in the event of a controller failure. Reduced system cost may also be beneficial, such as by adding new field devices based on small incremental investments rather than a heavy PLC burden.

[0293] Figure 31 The diagram illustrates a PLC-based industrial control system according to an example.

[0294] The advantages of the Multilayer Field Device Bus (MLFDB) described in this article can be understood by comparing it to a simplified traditional deployment based on a Programmable Logic Control (PLC). The most common way to implement a control strategy is through the use of a PLC that integrates control functionality, IO interfaces, and network access into a single device, such as Figure 31 As shown. A single PLC can be highly scalable, allowing the user to plug in many IO modules to expand the number of field devices in the control system. Although PLCs have served the industrial control system market well for decades, there are some limitations to this approach. First, if the PLC becomes inoperable, access to field devices and control functions becomes inaccessible. For reliability, industries have addressed this issue by purchasing two PLCs and two for each field device. However, this redundancy approach is expensive from a size, money, and power perspective. Second, small incremental changes to the infrastructure can require a significant investment since new PLCs may be required. In Figure 31 In the example, there are y IO modules in each PLC, where y is a finite number. The value of y can be based on the PLC vendor / model and can be in the range of 5 to 100, for example.

[0295] Figure 32 The figure shows a multi-layer field device bus (MLFDB) according to an example.

[0296] The difference between MLFDB and traditional PLC-based deployment is that the control function is completely separated from the field device IO. Figure 32 The decoupling of control functions from IO enables an “any-to-any” relationship between the controller and IO, which is a key capability for improving system reliability. Figure 32Each control function can access data from any connected field device, and likewise, each control function can control any connected field device. This "any-to-any" relationship improves system reliability by enabling built-in control function failover. For example, assume that control function 1 reads data from field device 2 (a level sensor), performs calculations, and adjusts the output value to field device 1 (a pump). If the device hosting control function 1 fails, the processes of control function 1 can be executed on another device with access to the field device bus. This is possible because the field device remains accessible on the bus.

[0297] Figure 33 The diagram illustrates the IO converter functionality according to an example.

[0298] The field device bus includes an IO converter. The IO converter is a device that can be addressed individually and converts the field device IO to the protocol of the field device bus. Figure 32 As shown, there are physically small, high reliability IO converters attached directly to each field device. The number of IO converters can range from 1 to n, where n is constrained by the physical environment of operation. Figure 33 A high-level view of the IO converter functionality is shown in a stacked view.

[0299] The IO converter is responsible for the following functions:

[0300] Electrical Interface to Field Devices: The interface from the IO converter to the field device can be any of 4-20mA analog input / analog output, 24VDC digital IO, serial interface, or Ethernet-based protocols. The design implementation of this interface determines the SKU of the IO converter. For example, IO converter SKU 1 can be a 4-20mA analog output or analog input. SKU 2 can be a discrete output for a high-current relay.

[0301] Field Device Protocol: This function encodes / decodes commands, controls, and data into the appropriate format required to communicate with downstream field devices. For example, assume the downstream field device is a Modbus slave. This function encodes a READ request that conforms to the Modbus protocol and sends it to the field device.

[0302] Abstraction: The abstraction function converts field device-specific commands and data into a human-readable format (defined by the data model). For example, suppose an IO converter is connected to a pump that communicates via a 4-20mA analog interface, and the control system wants to set the flow rate to 10 GPM. This function will convert the 10 GPM request into a corresponding current setpoint value in milliamps. Conversely, when data comes from the field device in a field device-specific format,

[0303] Information Modeling: This functional model models data according to a schema (eg, Haystack) defined by the system operator.

[0304] The field device bus protocol layer may conform to an industrial protocol for transmitting modeling data, for example, Data Distribution Service (DDS), QPCUA, or Profinet protocols.

[0305] Electrical interface to the field device bus. The electrical interface can include Ethernet, PCI, Profinet, proprietary bus technology, etc.

[0306] Provisioning functionality: The layer is a discovery layer that detects the identity of downstream field devices. The detection service may be built into the native field device protocol (e.g., HART), or it may need to be added as an additional discovery service. Either way, the provisioning layer represents the identity of the downstream connected field device.

[0307] The operational mode and resource status layer is responsible for reporting health and status to the orchestration system. Health and status data includes local resource utilization, workload status, unique module attributes, and operational mode.

[0308] Examples of local resource utilization may be CPU loading for reliable operation, memory utilization, storage, page misses, faults, etc.

[0309] The workload state will capture the status and health of running processes, and a crashed process can trigger an alert that the orchestration system can use to initiate failover conditions.

[0310] Unique module attributes consist of artifacts such as the IO converter's unique identifier (which can be hardware-based), IP address, MAC address, or certificates.

[0311] The operating mode refers to the role of an IO converter in a redundant system. For example, an IO converter can be placed in hot standby mode or in primary mode. Alternatively, an IO converter can be placed in a mode that electrically isolates it from a field device, such as by physically connecting a peer IO converter to the field device.

[0312] Agent: The agent residing on the IO converter arranges the configuration parameters for various IO converter functions.

[0313] Figure 32 or Figure 33The field device bus shown is not specific to a bus technology and can be instantiated with many different technologies. For example, Ethernet can be used in light of the increasing ubiquity of Ethernet-based devices in the industrial control system space. An advantage of an Ethernet-based field device abstraction bus is the increased accessibility of field devices to a wider range of systems. However, in order to maintain reliable and deterministic capabilities, an Ethernet-based field device abstraction bus may need to integrate time-sensitive networking (TSN). The integration of TSN can enable an Ethernet-based field device abstraction bus to match the reliability and timeliness of a Profinet-based or Ethercat-based system.

[0314] Figure 34 The diagram illustrates IO converter redundancy according to an example.

[0315] Multiple layers of redundancy can be used to address the situation where an IO converter directly connected to a field device fails. To mitigate this situation, Figure 34 As shown, multiple IO converters are added to the field device bus and physically connected to a single field device (in a multi-drop configuration). Each IO converter has 1..x switching outputs, of which only one can be actively driven at a time. This enables IO converter redundancy, which is controlled by an IO converter mode controller. The orchestration system can monitor the health and status of each IO converter, switching outputs on / off accordingly. The IO converter mode controller can change which IO converter controls which field device.

[0316] Figures 35A-35B Flowcharts 3500A-3500B are shown of methods for implementing an MLFDB according to an example.

[0317] Flowchart 3500A includes an operation 3510 for receiving data from a field device (e.g., a sensor) at an IO converter. Flowchart 3500A includes an operation 3520 for converting the data from the field device according to a field device bus protocol. Flowchart 3500A includes an operation 3530 for sending the converted data to a field device abstraction bus. Flowchart 3500A includes an operation 3540 for receiving a control signal from a control device via the field device abstraction bus. Flowchart 3500A includes an operation 3550 for sending an electrical signal to the field device based on the control signal.

[0318] Flowchart 3500B includes an operation 3560 for receiving data from a plurality of field devices (e.g., sensors) at a sensor bus via a plurality of corresponding IO converters. Flowchart 3500B includes an operation 3562 for sending the data to one or more control functions. Flowchart 3500B includes an operation 3564 for receiving one or more control signals from the one or more control functions based on the data. Flowchart 3500B includes an operation 3566 for sending the one or more control signals to corresponding IO converters in the plurality of IO converters. Flowchart 3500B includes an optional operation 3568 for receiving information from an IO converter mode controller. Flowchart 3500B includes an optional operation 3570 for facilitating the assignment of IO converters to field devices based on the information received from the IO converter mode controller.

[0319] Dynamic Alarms in Industrial Systems

[0320] Industrial control systems in supervisory mode rely heavily on alarms to protect machine operations. In many cases, these alarms are created based on human knowledge and understanding of the system. As a result, the alarms are suboptimal. A typical system will be enabled by a control engineer with numerous alarms for any condition considered suboptimal or harmful to the system. For example, an alarm might be created for voltage values ​​that fall below or rise above a certain threshold. Alarms in a control system are typically created for one of three reasons:

[0321] Safety (people and environment)

[0322] Equipment integrity

[0323] Quality Control

[0324] However, one of the problems in alarm management is that these alarms are generated by humans and are often prone to alarming about things that may not be important. Worse, alarms are often redundant because the same physical event may generate multiple alarms; this is often called alarm flooding.

[0325] Current alarm systems rely heavily on human generation and input. As a result, they tend to suffer from over-dispatching. This is easily seen in alarm floods, where large numbers of alarms are generated during a breakdown and often distract human attention.

[0326] Existing solutions tend to over-generate alarms, which is dangerous and can lead to alarm fatigue. This alarm generation is the result of excessive false positives used to detect situations where there are problems in the control system. If alarms result in a large number of events, it can also over-complicate analytics designed to use these events for other applications, such as anomaly detection.

[0327] The systems and methods described herein use intelligent machine learning to manage alerts. The systems and methods described herein can:

[0328] Characterize data to detect anomalies that could trigger alerts;

[0329] Clustering alerts using data similarity or common causal relationships so they are presented as a single bundle to combat alert flooding and fatigue; or

[0330] Understand how humans respond to alerts so you can automate those actions in the future.

[0331] Figure 36 An example of a process with generated alerts is illustrated according to an example.

[0332] In industrial systems, data is often generated by different modules and sensors. This data is the basis for alarm generation. In its most basic form, an alarm is generated based on a condition such as sensor data exceeding a threshold. For example, if a physical process is connected to a power meter, the control engineer can know that if the equipment (collectively) consumes more power than the circuit can withstand, an alarm can be issued to require human intervention. Alarms can go through multiple levels of escalation. For example, initially, an alarm is issued, but if power consumption continues to rise, additional alarms are generated and power is shut off from the system. The latter situation may be undesirable because it may incur costs associated with lost productivity and man-hours to restore the process to an operational state.

[0333] Alarms can have a cascading effect. For example, when one process shuts down, the next process along the factory line will stop, which may in turn trigger one or more alarms. Operators may suddenly find themselves in a deluge of alarms. Determining which alarm they should respond to and how can often be tiring and require further analysis and expertise.

[0334] The systems and methods described herein use machine learning to dispatch alerts, cluster alerts, or propose response actions.

[0335] Figure 36 An example physical process with an alarm generated by the system is shown.

[0336] These alerts can include one or more of the following data fields:

[0337] Alert Type

[0338] The physical process that generated the alert

[0339] Alarm Criticality

[0340] Timestamp

[0341] Possible signs or causes of alarms

[0342] User(s) marked as alert recipients

[0343] Possible actions to reset or resolve the alarm

[0344] Data can be sent to a central location, which can then be routed to an HMI screen, a user’s mobile device, or a repository for analysis and archiving.

[0345] In the systems and methods described herein, users can create alerts. These alert configurations can then be saved and analyzed. Based on the data, context, and alert configuration, additional alert suggestions can be presented to the user. For example, if an alert is created for an electricity meter with metadata indicating that the meter is on the factory floor, the metadata indicates that the meter is on the factory floor.

[0346] The similarity of other devices to these created alerts and their corresponding physical devices is analyzed using the plant’s metadata (and information model). In addition, the type of data generated by these devices and the flows that can trigger alerts are fed into the similarity module.

[0347] Figure 37 The figure illustrates a dynamic smart alert according to an example.

[0348] Figure 37 The system includes a data profiler, which can be referred to as a data signature manager. This module can use machine learning to determine flow similarities. Some of these similarities can be based on individual flows or as correlations between flows. For example, a level sensor flow can be compared to another level sensor flow generated by a similar physical process. For example, the physical processes can be considered similar based on the following:

[0349] Metadata about physical processes;

[0350] the number and types of flows associated with the same physical process;

[0351] Cross-correlations between different streams of the same physical process; or

[0352] Similarity in type and frequency of flows from different processes.

[0353] For example, when a first physical process has 20 flows and 3 of them are liquid levels and 2 are liquid flows, and a second physical process has 21 flows and 3 of them are liquid levels and 2 are liquid flows, these 2 flows can have a high ranking in similarity (same number of liquid levels and liquid flows, only differing by 1 flow).

[0354] The data signature manager feeds its output to the dynamic intelligent alert system, which acts as a classification unit to identify potential processes that need to adjust their alerts based on existing alerts from similar systems.

[0355] The dynamic intelligent alarm system can be composed of the following three components:

[0356] Alert Generator: In this module, a number of pre-alerts are pre-loaded or created by default. These can be created explicitly by humans or based on requirements. For example, the power consumption on a certain circuit should not exceed a certain threshold. This module is responsible for generating, editing, or deleting alerts. This module uses the output of data similarity to decide whether to create an alert or suggest an alert. It can create a score for the need for a specific alert. In the case of very high scores, the alert generator can automatically create an alert. However, if the score is moderate, the module can, for example, request input from a human operator / expert before creating such an alert. The alert generator's job is also to label alerts as similar, related, or independent. When multiple alerts are generated, the next module will use this label.

[0357] Alert Management / Clustering: This module tracks associations between alerts. It can use labels created by alert generators. It can also extend labels by monitoring data from actual alerts. Alert Management can monitor different alert outputs to detect correlations or sequences of events. It can run both types of analytics on the data.

[0358] Correlation can determine that two events are highly related and can be clustered together. For example, a specific physical process may have five different alarms that sound warnings for different events. However, if the system is down due to a serious fault, all alarms may be activated simultaneously or within a short period of time. These events are then highly correlated and can be clustered to minimize alarm flooding and fatigue. The module can also use metadata about the alarms and the systems they cover to create meaningful reasons for clustering components. Using the same example above, these five different alarms can have metadata about the physical processes associated with them. Therefore, these five can be collapsed into "Level Tank Processing in Area 3 West Building 2." Additionally, clustering can be meaningfully interpreted using data from the alarms themselves.

[0359] Figure 36Shows an example of what an alert might contain. The data can be aggregated and the result can be displayed as "Level: Critical, Reason: Power Too High". This clustering can use techniques from Natural Language Processing (NLP) to create these meaningful descriptions. Many of these descriptions can be human generated and may differ slightly when describing the same type of fault, and NLP can be used to align or group slightly different faults. In the example, if alert X notices that it will be forwarded to system Y, and then system Y causes the user to select a reset, then it may also be reset only at alert X. Initially, Figure 37 The system can proceed slowly, indicating a reset after X, and then over time, resetting without being asked.

[0360] The module can also model alarm sequences as state machines. For example, the module might notice that when process 1 fails, the probability of failure reported in process 2 is very high. Similar techniques can be used for predictive maintenance, where a sequence of events is modeled using a state machine and probabilities are assigned to the edges representing the transitions. This allows the algorithm to predict that if the system lands in state S1 and the transition between S1 and S2 has a high-probability edge, state S2 is likely to occur. This feature can allow the module to predict that another set of alarms is about to be triggered and potentially notify the user in advance. The relationship between alarms can be displayed before or after an event occurs.

[0361] Alarm Output Manager: This module is used to enable the system to enter autonomous operation. Initially, this module may have no policies or only some simple transcription policies. Then, as alarms are generated and processed, it can monitor user actions. If a set of alarms tends to be ignored, the module can learn over time that these alarms are meaningless and can be assigned a lower priority or even deleted. In some examples, deletion may not occur without human consent. Additionally, the module can monitor and record other events. For example, when an alarm is issued, a human operator can try several action options. These can include changing parameter configurations, resetting modules, and restarting parts of the system. For example, this sequence can be avoided by using the Alarm Output Manager to further characterize the specific characteristics of the fault, determining that a system restart is indeed necessary, or using a simple module reset. Alternatively, as the system's confidence increases, it can independently take these actions. In some examples, options may initially be presented to the human operator as recommendations.

[0362] Figure 38The figure illustrates a flowchart of a method for dynamic alarm control according to an example. Flowchart 3800 includes an operation 3810 for storing information related to a plurality of alarms of an industrial control system. Flowchart 3800 includes an operation 3820 for analyzing data, context, and alarm configurations for the plurality of alarms from the information. Flowchart 3800 includes an operation 3830 for suggesting changes to one or more of the plurality of alarms or suggesting a new alarm. Flowchart 3800 may include an operation 3840 for determining alarm flow similarities from the information. Flowchart 3800 may include an operation 3850 for detecting alarm events at two or more alarms. Flowchart 3800 may include an operation 3860 for preventing the two or more alarms from being issued. Flowchart 3800 may include an operation 3870 for generating clustered alarms for the two or more alarms that were prevented from being issued.

[0363] Autonomous ensemble methods for learning with closed-loop control operations

[0364] The integration of autonomous learning methods into practical, bounded implementations continues to grow across the industrial landscape, with significant progress being made in the robotics and autonomous driving spaces. As IT-OT convergence continues to materialize and enable greater modular system flexibility, the forward-thinking autonomous application of these developing technologies and methods will find its way into the broader continuous and discrete manufacturing industries. The ability to autonomously identify new models with proven value for mission-critical operations, and to autonomously deploy proven capabilities and "close the loop" with confidence, will unlock new levels of efficiency, cost savings, and bottom-line value for manufacturers operating in control systems compliant with IEC 61131-3, IEC 61499, and higher ISA levels (L1-L3).

[0365] The integration of traditional closed-loop control systems with autonomous learning technologies requires the creation of new resilient solution architecture approaches to support autonomous workflows. These approaches will also inherently generate new autonomously developed closed-loop control solution architecture proposals, which will need to be evaluated for feasible implementations that fit within the defined reference architecture boundaries for both continuous and discrete manufacturing operations. Such automatically "closed-loop" autonomous systems can support real-time policy evaluation for safety, quality, constraint identification, implementation feasibility, value scoring, automated monitoring, and system management integration for viable mission-critical system deployments. Autonomy can be extended beyond pure software integration and can be integrated with all aspects of end-to-end system deployment, including hardware selection across compute, storage, and networking assets. For any specific new control application created, real-time coordination and verification across multiple subsystem domains may be required to ensure safe and bounded operation of closed-loop solutions for autonomous deployments.

[0366] This paper proposes a rigorously ordered policy framework and a set of methods for managing the autonomous creation of new closed-loop workloads in mission-critical environments through the following eight processes:

[0367] Assessment of the quality and sensitivity of the new algorithm relative to the process;

[0368] Automated establishment of operational constraint boundaries;

[0369] Automated safety assessment of new algorithms relative to existing processes;

[0370] assessing the value of automation for a wider range of processes;

[0371] Automated system assessment for deployment feasibility in controlled environments;

[0372] Physical deployment and monitoring of new application control policies;

[0373] Integration into lifecycle management systems; and

[0374] Integrated into scrap handling.

[0375] The order of operations in this 8-step process can be changed. For example, the safety assessment can be performed after the value assessment.

[0376] Typical automation systems, while very advanced in their implementation of control strategies, are inherently locked into legacy system deployments where this type of system resilience often does not exist. New systems can have new levels of flexibility and resilience that reflect the system advancements found in IT systems, and none of the current IT or OT systems possess this level of autonomous intelligence.

[0377] Typically, previous solutions implemented in the distributed control system design space do not allow for any level of autonomous creation of new control strategies, with subsequent implementation and debugging (closing the loop) without a high degree of engineering oversight. Furthermore, current control strategy design is not autonomous and requires a high degree of engineering. Control implementation is also a highly resource-intensive engineering activity. Control debugging activities, where the loop is closed and the algorithm is tuned, are also manual and highly engineered processes. Currently, accomplishing any of these tasks autonomously is unheard of in practice.

[0378] Prior solutions may not have leveraged the ability to create automated general safety assessments of newly created algorithms relative to existing processes. Prior solutions may not have leveraged the ability to create automated quality and sensitivity assessments of new algorithms relative to existing processes. Prior solutions may not have leveraged the ability to create automated establishment of operational constraint boundaries. Prior solutions may not have leveraged the ability to create automated system assessments for deployment feasibility into a control environment. Prior solutions may not have leveraged the ability to create automated value assessments for the broader process based on available data.

[0379] Previous solutions may not have leveraged the ability to create automated physical deployment and monitoring of new control applications. Previous solutions may not have leveraged the ability to create automated integrations into existing standardized lifecycle management systems. Previous solutions may have been locked into application and device-specific implementations where dynamic workload modification and portability were impossible. Previous solutions may have been tightly coupled to hardware and software. In most cases, previous solutions may have been prohibitively expensive. Previous solutions may have required custom hardware with custom interrupt management. Previous solutions may not have included dynamic discovery, simulation, and optimization that predict value events as part of the rule set used for the decision tree.

[0380] A strictly sequential policy framework and a set of methodologies are proposed here for managing the autonomous creation of new closed-loop workloads in mission-critical environments through these eight steps (which can occur in the following order, or in other orders, or with some steps occurring in the same timeframe or overlapping):

[0381] Assessment of the quality and sensitivity of the new algorithm relative to the process;

[0382] Automated establishment of operational constraint boundaries;

[0383] Automated safety assessment of new algorithms relative to existing processes;

[0384] assessing the value of automation for a wider range of processes;

[0385] Automated system assessment for deployment feasibility in controlled environments;

[0386] Physical deployment and monitoring of new application control policies;

[0387] Integration into lifecycle management systems; and

[0388] Integrated into scrap handling.

[0389] The systems and techniques described herein provide distributed application control policy-level capabilities for:

[0390] The overall learning system operation under control can be integrated with the existing control system hierarchy.

[0391] Under control, automated safety assessments are performed for new algorithms relative to existing operational processes.

[0392] The quality and sensitivity of new algorithms relative to existing physical processes are evaluated under control.

[0393] Automatically establish operational constraint boundaries for the system under control.

[0394] Enables automated system assessment of deployment feasibility of autonomously created applications for controlled environments.

[0395] Automated value assessment for the wider processes under control is achieved to ensure the positive economic impact of autonomously created control algorithms.

[0396] A capability is implemented for autonomously physically deploying new autonomously created control applications and creating new monitoring for the new autonomously created control applications.

[0397] Enables autonomous integration into standard lifecycle management systems.

[0398] Enables integration into scrap processing through continuous ROI monitoring.

[0399] Continue to advance and leverage lower-cost commodity hardware and software to achieve better system performance at the control strategy level.

[0400] Enable many maintenance tasks to occur autonomously, where autonomous functionality is designed into the automation configuration.

[0401] Figure 39 The figure shows the autonomous control-learning integration flow chart.

[0402] A strictly sequential policy framework and a set of methodologies are proposed here for managing the autonomous creation of new closed-loop workloads in mission-critical environments through these eight steps (which can occur in the following order, or in other orders, or with some steps occurring in the same timeframe or overlapping):

[0403] Assessment of the quality and sensitivity of the new algorithm relative to the process;

[0404] Automated establishment of operational constraint boundaries;

[0405] Automated safety assessment of new algorithms relative to existing processes;

[0406] assessing the value of automation for a wider range of processes;

[0407] Automated system assessment for deployment feasibility in controlled environments;

[0408] Physical deployment and monitoring of new application control policies;

[0409] Integration into lifecycle management systems; and

[0410] Integrated into scrap handling.

[0411] The interaction of these eight sequential processes is shown below and Figure 39 Each process is described in more detail in , with iterative feedback analysis as the foundation to support continuous 24 / 7 mission-critical operations.

[0412] A. Creation of New Learning Algorithms

[0413] The process can start with the creation of a new learning algorithm. The autonomous process used will have system-wide access to all data associated with system resources, physical processes, and control system parameters, including basic ISA Level 1 control (IEC61131-3 / IEC61499-3 function blocks, binary, ladder logic, PID, etc.), constraint and supervisory control (L2 / L3), multivariable model predictive control (L3), production scheduling (L3), and planning system access (enterprise). The learning system scope can include unconventional system access associated with finance and accounting, contract management, and general enterprise or supply chain operations. Although the algorithms created using significant correlations can cover a wide range of algorithm families, including simple small data-oriented mathematical solutions (sum, division, multiplication, PID, statistics, etc.), autonomous model development based on first principles, autonomous model development based on experience, big data analysis, machine learning, and deep learning algorithms, etc., this disclosure does not describe the complete environment of the algorithms that can be generated. From the perspective of data science, such algorithms can be open.

[0414] As the analysis moves sequentially from Step A to Step I, pass and fail tests are autonomously created and executed to evaluate and validate the new autonomous learning and control loops created. An iterative process is employed to support real-time pass / fail analysis. Some steps may be executed out of sequence, concurrently, etc.

[0415] B. Quality and Sensitivity Assessment

[0416] Once a significant model is autonomously discovered in step A, step B is called to form an initial quality and sensitivity assessment of the created algorithm. The autonomous quality and sensitivity assessment relies on a process model (digital twin) of the latest real-time simulation of the process. The simulated process scope can be a subset of the entire process, which can include items as small as a valve or pump, or cover the complete process unit under control (refinery crude unit, reactor, or a wider area of ​​the plant). This general quality and sensitivity assessment takes the model created in step A and overlays it on the simulated physical process and control algorithm actively used in the distributed control system. The quality and sensitivity assessment is then performed for each independent process variable by generating input signals (PRBS, Schroeder waves, etc.) for the new model, and tracking the impact on the relevant process variables over time for the active control strategies deployed in the simulated process and system. For the quality assessment profile that considers the sensitivity of the new model to the simulated process operation, the process output results are measured both absolutely and statistically.

[0417] Table 1: Architectural subsystem evaluation for quality and sensitivity

[0418]

[0419]

[0420] Table 2 Quality and sensitivity evaluation

[0421]

[0422] If the test passes, the model evaluation moves to step C. If the test fails, the results are sent back to the learning system for re-evaluation.

[0423] C. Constraint Boundary Identification

[0424] The results of step B are used to set constraint boundaries for the new model created, which encompass and enforce the process-wide operational safety, quality, and sensitivity criteria for the new algorithm created. The identified new constraint boundaries are then run through simulation (e.g., adding noise, perturbations, and seeing how the system reacts using the new model) and the results are compared to the newly generated constraint profile.

[0425] Table 3 Evaluation of architectural subsystems for constraint boundary identification

[0426]

[0427] Table 4 Process constraint boundary identification

[0428]

[0429]

[0430] If the test passes the evaluation of the process simulation and associated existing and new control models, the evaluation is allowed to move to step D. If the test fails any of the criteria described above, the results are sent back to the learning system for re-evaluation.

[0431] D. Security Assessment

[0432] Once a new set of constraints has been identified in step C, step D is invoked to form an initial safety assessment of the created algorithm. Step D encompasses a safety assessment of the relative impact of the new learning algorithm on the physical process used for closed-loop operation using the new learning algorithm. Here, model quality is evaluated over a range of conditions by introducing noise into the model I / O, building on the statistical quality established during model creation in step A.

[0433] The autonomous safety assessment relies on a process model (digital twin) of the latest real-time simulation of the process. The scope of the simulated process can be a subset of the entire process, which can include items as small as a valve or pump, or cover the complete process unit under control (a refinery crude unit, a reactor, or a wider area of ​​the plant). This general safety assessment takes the model created in step A and overlays it on the simulated physical process and the control algorithms actively used in the distributed control system, and overlays the new constraints identified by the control system as described in step C. The safety assessment is then performed for each independent process variable by generating input signals (PRBS, Schroeder waves, etc.) for the new model, and tracking the impact on the relevant process variables over time for the simulated process, the new constraints deployed in the system, and the active control strategies. The results are then compared to the safety metrics established for the process that are used to determine a pass / fail score.

[0434] There is a wide range of potential safety checks that can be used and are not covered here, but they may generally appear as key process constraints on flow, pressure, temperature, rpm, or other key variables that may not be exceeded for normal safe operation. This analysis is not to be confused with certified functionally safe systems, although the effects on variables associated with these systems are considered to be within the scope of general safety analysis and may in practice be included in process and control simulations.

[0435] Table 5: Evaluation of architecture subsystems for security

[0436]

[0437] Table 6 General process safety assessment (uncertified FuSA)

[0438]

[0439]

[0440] The results are measured against the safety profile for the device or process being controlled. If the algorithm passes all safety checks for the device or process flow for the manufacturing operation, the validation process is allowed to move to step E. If any safety check fails, the results are returned to the autonomous learning algorithm block to re-evaluate and recreate the saliency algorithm.

[0441] E. Value Assessment

[0442] Value assessment is used to autonomously evaluate the impact of new models on local process segments and the broader end-to-end manufacturing process. Leveraging boundary constraints identified for the simulated process (digital twin) and control system, the impact of new learning algorithms on the enterprise bottom line is automatically assessed within the context of closed-loop performance by replaying historical digital twin simulation results using the new control strategy. Results are compared to baseline performance using various value metrics. Examples are shown in Table 7 below.

[0443] Table 7 Example Value Assessment Criteria

[0444]

[0445]

[0446] If the value assessment meets the specified RQI and NPV criteria specified by the operation, the test passes and the evaluation proceeds to step F. If the test fails, the results are sent back to the learning system for further evaluation.

[0447] F. Deployment Feasibility

[0448] Deployment feasibility is measured by the system's ability to deploy new workloads and integrate algorithms into the existing control structure of the distributed control system. This process involves a rigorous real-time evaluation of the following areas:

[0449] Table 8 Deployment feasibility subsystem evaluation

[0450]

[0451]

[0452] If the deployment test passes, the deployment is then tested by actually deploying to a digital twin simulation system where the training is automatically scheduled using operations.

[0453] Automated training simulator deployment and process scheduling:

[0454] i. Generate automated training documentation for the new control loop and send it to operations for review.

[0455] ii. Training plans are established and completed based on operations.

[0456] G. Physical Deployment and Monitoring

[0457] After the physical deployment and monitoring have been tested in step F using the new constraints identified for the control system, training and exit through operations are completed under the new control strategy ready for deployment. The physical implementation is managed by the system orchestrator, which specifies the input and output configuration of the new function blocks and modifications to the old function blocks. The steps for a feasible autonomous deployment are as follows:

[0458] Online deployment to operation:

[0459] 1. Programmatically, all control may be automatically relinquished to its lowest permissible autonomous stable loop configuration, as pre-specified for the operation of autonomous system implementation and commissioning of new control system features.

[0460] 2. Complete the physical deployment of the new control and learning model(s) within the defined system constraints and monitoring configuration using available system resources (compute, storage, networking, etc.).

[0461] 3. Automated commissioning occurs when a new loop automatically begins running in “warm mode,” where live I / O is fed into the new control loop and the new control actions of the independent variables are analyzed over a specified period of time to verify that the behavior is as expected.

[0462] 4. Once online validation testing is complete, the loop is closed for the new algorithm and the new output is written to a downstream set point that drives the mission-critical process with notifications sent to operations.

[0463] If the autonomous deployment is successful for all four steps described above, the system proceeds to automatically register with the lifecycle service. If the autonomous physical deployment of the system fails in any of the four steps described above, the system is returned to its previous configuration, the results are sent back to the learning system, and the action is notified.

[0464] H. Lifecycle Integration

[0465] Automatic enrollment consists of a new control and learning cycle for the system deployed for normal operations.

[0466] l. Generate, test, and deploy automation scripts to register new control applications into the lifecycle management system.

[0467] 2. Feedback loop targets Figure 39 Automated metrics for quality, constraints, security, value, deployment, and lifecycle performance are shown to continuously monitor new control applications.

[0468] A degradation in any of the monitored metrics may send the currently running control strategy back to the control learning evaluation block, or result in a change in the limit specifications, or a change in tuning parameters, a change in deployment, etc.

[0469] I. Scrap

[0470] Using feedback (see Figure 39 ), continuous checking of economic value assessment metrics can drive automated assessment of the operational value of an enterprise. A drop in value below a defined threshold triggers the scrapping process.

[0471] While scrapping can be automated, the option to revert to manual review will be required and may result in automatic deactivation or require manual removal depending on the complexity of the automation.

[0472] Figure 40 The figure illustrates a flowchart of a method for managing the autonomous creation of new algorithms for an industrial control system, according to an example. Flowchart 4000 includes an operation 4010 for managing the autonomous creation of a new closed-loop workload algorithm. Flowchart 4000 includes an operation 4020 for performing a quality and sensitivity assessment of the new algorithm relative to the process. Flowchart 4000 includes an operation 4030 for autonomously establishing operational constraint boundaries. Flowchart 4000 includes an operation 4040 for autonomously assessing the safety of the new algorithm relative to the existing process. Flowchart 4000 includes an operation 4050 for autonomously assessing the value of the new algorithm for the broader process. Flowchart 4000 includes an operation 4060 for autonomously assessing the feasibility of deploying the system in a control environment. Flowchart 4000 includes an operation 4070 for physically deploying and monitoring the new application control policy. Flowchart 4000 includes an operation 4080 for integrating the new algorithm into a lifecycle management system. Flowchart 4000 includes an operation 4090 for integrating the new algorithm into the decommissioning process.

[0473] Scalable edge computing in distributed control environments

[0474] Current solutions require end users to estimate the amount of compute required and add additional computing capacity to future-proof their deployments. This approach wastes money, electricity, and heat. It also runs the risk of over-provisioning compute that becomes legacy technology before it's actually needed.

[0475] The technology discussed in this article allows high-performance CPUs in edge control nodes of industrial control systems to be activated from an initially dormant or inactive state by a centralized orchestration system that understands the CPU performance requirements of the industrial system's control strategy. Because each edge control node is initially sold as a low-cost, low-performance device, the initial customer investment is low. Only the required compute (right-sized compute) is purchased and provisioned, optimizing monetary investment, thermal footprint, and power consumption. This solution provides a scalable compute footprint within the control system.

[0476] Figure 41 The diagram illustrates an industrial control system (ICS) ring topology network 4102 .

[0477] Industrial control systems typically consist of a programmable logic controller 4104, remote IO (RIO) (e.g., 4106), and field devices (e.g., 4108). A typical deployment may include a ring of remote IO units controlled by a PLC 4104. It is common to lock the IO and field computing into the PLC 4104 (e.g., in Figure 41 middle).

[0478] Figure 42 The diagram illustrates an edge control topology network. The edge control topology network includes an orchestration server 4202 (e.g., as described above for orchestration 920), a bridge 4204, multiple edge control nodes (e.g., ECN 4206), and one or more field devices (e.g., 4208). Orchestration server 4202 is used to provision, control, or orchestrate actions at ECN devices (e.g., 4206), which are interconnected, for example, in a ring network and connected to orchestration server 4202 via bridge 4204.

[0479] One way SDIS improves system functionality is by distributing control functions across ICSs. Orchestration server 4202 can be used to control edge control nodes 4206, which include the option to perform both IO and compute on a single device and use orchestration services to assign workloads to the best available resources.

[0480] Typically, a ring of edge control nodes (ECNs) may be deployed in thermally constrained environments, such as cabinets with zero airflow or unregulated temperatures. In an example, a single cabinet may have up to 96 I / Os, which means up to 96 ECNs. This may prohibit each ECN from including both I / O and high-performance computing, as the high-performance computing equipment would generate excessive heat and raise the ambient temperature above the ECN's safe operating level. Additionally, when the computational requirements of the control system are not high, it may not be necessary to have a high-performance processor at every ECN. Therefore, the systems and techniques described herein provide the ability to install only the computing resources required to execute the control strategy without exceeding cost and power targets, while still allowing changes in each ECN. Therefore, in an example, not every ECN has a high-performance processor or high control capabilities.

[0481] Figure 43 The diagram shows an edge control node (ECN) block diagram 4302. In the example, the following techniques are introduced by Figure 43 The computation shown scales the ECN to provide a "right-sized" supply for the computational problem.

[0482] The main components of ECN 4302 may be a system on chip 4304 having a higher performance computing (e.g., CPU) 4306 and a microprocessor (MCU) 4308 for low performance computing. MCU 4308 may be used to convert IO data from IO subsystem 4312 to network components 4310, such as Ethernet TSN-based middleware, such as OPC UA publish / subscribe or DBS. ECN 4302 may be delivered to customers with high performance CPU 4306 in an inactive state. For example, high performance CPU 4306 may not be accessible in an inactive state, such as until a special "activation signal" is sent to high performance CPU 4306 from, for example, an orchestrator (e.g., the orchestrator may send a signal to MCU 4308 to activate CPU 4306).

[0483] ECN 4302 may initially be installed as a low-cost, low-power device for IO conversion using MCU 4308. For example, high-performance CPU 4306 is initially disabled, and initially, ECN 4302 includes SoC 4304 and IO subsystem 4312 activated without high control capabilities. In this example, high-performance processor 4306 may be inactive, and ECN 4302 initially only allows IO conversion.

[0484] Figure 44 The figure shows a ring topology based on ECN. Figure 44 Shown is how scalable computing ECN can be adapted to the classic ring topology. Figure 44Also shown is the initial state of the deployment, where all high-performance CPUs are disabled. Figure 44 As shown, each ECN has the ability to convert IO to a data bus standard, but has no real ability to perform control functions.

[0485] In an example, after deployment, orchestration server 4202 can determine how many high-performance CPUs are needed and then send code to activate one or more CPUs using the corresponding MCU at a specific ECN. Orchestration server 4202 can provide a cost / benefit analysis as part of the scheduling function performed by orchestration server 4202. In an example, a fee can be charged to activate CPU 4306, for example, based on a schedule such as a monthly or annual license. CPU 4306 can be activated or deactivated as needed (e.g., as determined by the orchestrator or the user). A limited license can be less expensive than a full deployment. In another example, once activated, CPU 4306 can remain activated indefinitely (e.g., permanently activated for a one-time fee).

[0486] In an example, deactivating CPU 4306 can reduce heat output. This can be controlled separately from any power scheduling. For example, once activated, CPU 4306 can be deactivated or moved to a low-power state to save heat output (even in examples where CPU 4306 is permanently activated). CPU 4306 can execute control instructions in a high-power state and move to a low-power state when execution is complete.

[0487] In an example, the activation code may be a special packet sent to the MCU 4308. The validity of the activation code may be evaluated by the MCU 4308, including determining how long the code is valid, etc. The MCU 4308 may send the activation signal directly to the CPU (e.g., after receiving a signal from the orchestrator).

[0488] When activating the CPU 4306 from an inactive state, the MCU 4308 may turn on the power rails, boot the CPU 4306, download the latest firmware, etc. In an example, the CPU 4306 may have a low or high power mode that may be activated or deactivated rather than turning off or on the CPU 4306. This example may be useful in situations where the CPU 4306 may be placed in a low power state rather than powered off to reduce heat output, such as when the CPU 4306 may need to be activated quickly.

[0489] In an example, a low power state can be achieved by providing an encrypted token obtained from the CPU manufacturer by the orchestrator 4202. These tokens can be sent to the CPU 4306 via the MCU 4308. For example, the tokens can be signed using a key known only to the CPU manufacturer and the CPU 4306 (e.g., a key burned into the CPU 4306 during manufacture), thereby allowing each token to be verified. Each token can be unique, thereby allowing the CPU 4306 to operate for a period of time.

[0490] In another example, the token is authenticated by the MCU 4308 using a secret known to the manufacturer and the MCU 4308. For example, as long as the MCU 4308 and the CPU 4306 are manufactured together in a single package in an SoC, this example can prevent a denial of service attack that is created by waking up the CPU 4306 to verify the token.

[0491] Figure 45 The diagram shows the data flow through an ECN-based ring topology. In this example, the orchestration system 4202 analyzes the control strategy to understand how much computing is needed to meet the computing requirements of the control strategy. Once the orchestration system generates the computing requirements, the end user can purchase the required amount of high-performance CPU activation codes from the ECN provider. The orchestration system 4202 sends the authenticated activation code to the designated ECN in the ECN array, which enables the computing resources. This process is described in detail in the following sections. Figure 45 Shown in.

[0492] The process of enabling compute doesn't have to be a one-time event. As control strategies grow in complexity and compute demands increase, end users can continue to purchase and activate more compute resources (or deactivate CPU resources when they are no longer needed). For example, the orchestrator can send a deactivation signal to the ECN to deactivate the CPU at that ECN. ECN providers can implement a time-based service model in which end users purchase activation licenses on a monthly or annual basis. This model also allows end users to let activation codes expire, allowing some compute resources to return to a low-power sleep state, thus saving recurring costs.

[0493] Figure 46AThe diagram illustrates a flowchart 4600A of a method for activating a CPU (e.g., of an ECN) according to an example. Flowchart 4600A includes an operation 4610A for determining, at an orchestration server, the computing requirements of an edge control node in an industrial control system (e.g., a ring deployment). Flowchart 4600A includes an operation 4620A for receiving an indication to activate the CPUs of one or more edge control nodes or to determine that one or more CPUs need to be activated. Flowchart 4600A includes an operation 4630A for sending an authenticated activation code to the edge control node having the CPU to be activated. In the example, operations 4610A-4630A (above) can be performed by the orchestration server, and operations 4640A-4670A (below) can be performed by the ECN. The method using flowchart 4600A can include performing operations 4610A-4630A or 4640A-4670A or both.

[0494] Flowchart 4600A includes an operation 4640A for receiving an authenticated activation code at an edge control node. Flowchart 4600A includes an operation 4650A for authenticating the code at the edge control node. Flowchart 4600A includes an operation 4660A for activating the CPU of the edge control node using an MCU (low-performance processor). Flowchart 4600A includes an optional operation 4670A for receiving an update from an orchestration server at the edge control node to deactivate the CPU or place the CPU in a low-power state. In an example, the ECN can be part of a ring network of an industrial control system.

[0495] Figure 46B 4600B illustrates a method for activating a CPU according to an example. The operations of flowchart 4600B may be performed by an orchestration server. The orchestration server may be communicatively coupled to a ring network of edge control nodes, such as via a bridge device.

[0496] Flowchart 4600B includes optional operation 4610B for determining the computational requirements of an edge control node in an industrial control system. In an example, the edge control node can be a node in a ring topology network, where a bridge device connects the network to an orchestration server.

[0497] Flowchart 4600B includes an operation 4620B for receiving data via a bridge connecting an orchestration server to an edge control node. The IO data may be converted at the edge control node's microcontroller (MCU) from data generated by the IO subsystem. This conversion may be to packets sent by an Ethernet switch on a system-on-chip (SoC) of the edge control node (which may also include an MCU). In another example, the data converted by the MCU may be data generated by the MCU itself, such as the power status of a field device or edge control node.

[0498] Flowchart 4600B includes operation 4630B for sending an authenticated activation code to an edge control node to activate an initially inactive CPU of the edge control node. In an example, the authenticated activation code is authenticated by the MCU before the CPU is activated.

[0499] Flowchart 4600B includes an operation 4640B for sending processing instructions to the CPU for execution.

[0500] Flowchart 4600B includes an optional operation 4650B for sending a deactivation code to the edge control node to deactivate the CPU of the edge control node.

[0501] The method may include an operation for determining computing requirements of a plurality of edge control nodes in an industrial control system including the edge control node. In an example, the CPU is activated based on an orchestration server determining that the CPU should be activated to satisfy a control policy for the industrial control system. In another example, the orchestration server may receive an instruction to activate the CPU of the edge control node among the plurality of edge control nodes.

[0502] Distributed dynamic architecture for applications and client-server frameworks

[0503] In an orchestrated system, an application is defined as a set of modules interconnected by a topology. These modules are deployed on different logical nodes. Each logical node can correspond to a physical node, but the mapping does not have to be 1:1. As long as resource requirements are met, multiple logical nodes can be mapped to a single physical node, or multiple modules can be deployed on the same physical environment.

[0504] As different modules are deployed, various errors, crashes, or restarts of modules or nodes may occur. In order to improve the resilience of the deployed application, redundancy can be used to improve availability. For example, a module can be deployed on two nodes (e.g., as a primary node and a backup node). When the primary node has an error or otherwise fails, the orchestrator can switch to the backup node to allow it to take over. However, the saved state of the closed modules is generally non-trivial. In the systems and technologies disclosed herein, the system includes a peer relationship between nodes on the same level in the application topology, which can act as automatic backup nodes or coordinate to generate backups. The use of peer coordination can allow the use of saved state, which can include monitoring the communication channel and redeploying the module on a different node in the event of a module or node failure or crash.

[0505] Current redundancy solutions are manually defined or created through redundancy. This results in high reliability, but also comes at a significant cost because it requires duplication of resources. Manual redundancy is often difficult to define and maintain. Strategies are often overly simplistic and require excessive resources. Furthermore, requiring a central orchestrator to identify redundant nodes or replace failed ones is both expensive and slow.

[0506] In an example, the technology described herein can create automatic redundant nodes based on the communication mode of the application of the module. For example, when the first module sends data to the second module, the node that controls the second module can become automatically redundant for the first module. The data generated by the first module is fed into the second module, allowing the first module to know what the input of the second module is. When the first module sends data to multiple modules instead of just to the second module, other problems may occur (or when the second module receives input from a module other than the first module). In these scenarios, it may be difficult to create redundancy on any of these leaf nodes. In contrast, a peer-to-peer network created by a set of nodes on the same layer can negotiate the status of redundant nodes. A network of these nodes can exchange redundant sets between the nodes themselves without major impact on the rest of the application.

[0507] Figure 47 The figure shows an example application connection diagram. In the example, the different modules that form the application can be connected in a way such as Figure 47 The example shown is configured to be arranged in the example shown. The connections show the flow of data between the different modules. The modules use communication channels that can operate in client / server or publish / subscribe mode to send data. In this example, when the orchestrator deploys the modules, the orchestrator can choose to deploy each module on a separate compute node or to deploy multiple modules on a single node. In this example, for simplicity, a single module is deployed on a single node. Other examples can provide redundancy options when multiple modules are on a failed node, or when a module has an error (for example, when another module on the node has no error).

[0508] In the example, module B on node 4720 is sending data to both module E on node 4740 and module D on node 4730. When module B fails, the following operations may be performed. The operations may be performed by peer nodes such as node 4710, node 4730, and node 4740. The execution may include detecting the failure as needed, redeploying module B on a replacement node (e.g., when node 4720 fails), rerouting inputs (e.g., from module A) or outputs (e.g., to module E or module D), and restoring the previous state of module B, which may be transferred to the replacement node.

[0509] exist Figure 47In the example shown, the neighbors of module B (e.g., module A, module D, and module E) can create a peer-to-peer network with the goal of taking over if module B fails (e.g., when node 4720 fails). In this example, the neighboring modules are positioned to recreate the state of module B because modules A, D, and E are in direct contact with the input and output channels of module B. These three neighboring modules can undergo a leader election algorithm or other technique for selecting a replacement node.

[0510] In an example, the executable file for module B can be deployed on one or more of the three nodes (e.g., 4710, 4730, or 4740), or one or more of the three nodes can manage where the redundant software resides. In an example, in the event of a failure of node 4720, one or more of the three nodes can manage routing inputs or outputs. In another example, data can be routed even if a failure is not detected (e.g., for redundancy purposes). Backing up module B using one of these techniques allows for a seamless switch to redundant nodes in the event of a failure because these nodes control where data flows. In an example, one or more redundant nodes can utilize software running shadow nodes as redundancy throughout the operating cycle.

[0511] exist Figure 47 In the example shown, module B has modules A, D, and E as neighbors. These four modules establish a neighborhood around B (e.g., a peer-to-peer network) and create a contingency plan for when module B fails. This plan may include using a leader election algorithm or other techniques to select a control node (e.g., node 4710 is elected as a redundant node with more resources to run module B, such as on the additional resources of node 4710). The control node or the selected replacement node may not be directly connected to the failed node 4720 and may store the redundant version of module B. When node 4720 fails, a redundant version of module B exists, and the redundant node can then seamlessly execute module B. For example, module A can create a channel to let module B know about the redundant node running the redundant version of module B. Module B and the redundant version can then be linked together, and in the event of a module B crash, module B can send status details to the redundant module to keep the redundant module informed.

[0512] Figure 48 The figure shows an example architecture diagram of an application with redundant nodes. Figure 48 In Figure 4, the three nodes (4810, 4830, and 4840) of master module A, module D, and module E form a peer-to-peer network. Module A is the leader of the network and manages master module B' on redundant node 4825. Module A can also route its output as input to both nodes 4820 and 4825. Figure 48In the example above, module B' is continuously computing outputs (eg, the same as module B) even though module B' is not connected to any modules.

[0513] With this arrangement, the application has its own resilience independent of the orchestrator 4805 (which can be used to set up the application or network configuration and then disconnected). The independence of the application can allow it to be completely disconnected from the orchestrator 4805 without sacrificing reliability.

[0514] In some examples, when the physical node of the master module is resource-constrained, it may not be feasible to have module B' run all the computations. However, to achieve full redundancy, one of the options described below can be implemented.

[0515] One option involves executing module B in a virtual machine. In this example, the system can make a copy of the virtual machine as soon as available resources allow, without compromising the operation of the rest of the application (e.g., by waiting for downtime or for additional resources on the node to become available). By doing so, the state of module B can be preserved (e.g., as an image of the virtual machine).

[0516] In another option, module B can support swapping, which allows module B to have an interface for submitting its internal parameters and state information to module B'. This redundant operation can be performed regularly to allow module B to preserve its state. The frequency of updates can depend on the size of module B and whether it can be updated while continuing to meet the requirements of the different modules and the overall application.

[0517] In this example, when module D is elected as the leader, module D can listen to all channels required by module B' to ensure that data is not lost (for example, the output from module A). This allows data to be forwarded to module B' when needed. Similarly, module D can set module B to listen to a channel (for example, the output from module A) without module D having to listen to that channel directly.

[0518] In some examples, the orchestrator or application developer may determine that a module is too critical to the application or is a single point of failure. In this case, more than one redundant module can be assigned to that module. For example, a network formed by three nodes can then create multiple redundant modules (e.g., module B' and module B", not shown). Each of these modules can have a different synchronization strategy to create diversity or increase resiliency.

[0519] Typically, applications do not exist in a walled garden, but rather are often connected to other applications. Similar to the techniques and systems described above, replacing modules with applications allows the system to provide redundancy at the micro or macro level. For example, Application I can connect to Application II and become the leader in creating redundancy and redundancy strategies (e.g., in the event of a failure of that application).

[0520] In the event of a cascading failure or major outage, creating such a strategy and allowing applications to have their own strategies can provide redundancy without unnecessary costs. Fully distributed systems are generally more difficult to manage, but have a higher degree of resilience due to the lack of a central authority that could become a single point of failure. Therefore, in this case, each application can have its own reliability strategy (policy, strategy). In an example, applications can be interconnected and apply their own macro reliability strategy. In an example, when two or more modules, nodes, or applications fail, the remaining modules, nodes, or applications can act as redundancy for the failure. For example, if two nodes fail, a single node can replace the two failed nodes, or two or more nodes can replace the two failed nodes.

[0521] When a system is under security attack, redundant applications or modules with macro or micro reliability strategies can provide protection. Multiple failures can be detected at the macro level, and accordingly, the strategy can be changed. For example, when a failure threatens to wipe out nearby applications, the deployed strategy can purposefully assign distant neighbors to be part of a community to save states, modules, or applications from the total failure. When in Figure 48 When considering security in the example of , module F or module C can join the network and be assigned a role. This role may not be a leader but a member of the community. In other words, module C may not expend too many resources managing module B'. Instead, module C may make a redundant copy of module B (for example, occasionally) but not instantiate it. This may sacrifice some seamless properties (for example, the state may be a little stale), but will provide an additional layer of guarantees and redundancy at minimal cost to the entire system. The same concept can be applied to applications so that if part of a local data center becomes unavailable, another data center in a different location can take over with slightly stale state and internal variable values, allowing operations to continue.

[0522] Figure 49AThe diagram illustrates a flowchart of a method for creating an automatic redundant module for an application on a redundant node based on the application's communication pattern, according to an example. Flowchart 4900A includes an operation 4910A for creating a peer-to-peer neighbor network. Flowchart 4900A includes an operation 4920A for presenting a redundant module on the redundant node, the redundant module corresponding to a module of the application on the node. Flowchart 4900A includes an operation 4930A for detecting a failure of the node for the module. Flowchart 4900A includes an operation 4940A for activating the redundant module on the redundant node by reconnecting the module's inputs and outputs to the redundant module. Flowchart 4900A includes an operation 4950A for restoring a previous state from the module and transferring it to the redundant module. Flowchart 4900A includes an operation 4960A for continuing execution of the module using the redundant module. Flowchart 4900A includes an operation 4970A for reporting a failure of the node.

[0523] Figure 49B 4900B is a flowchart illustrating a method for activating a CPU according to an example. The operations of flowchart 4900B may be performed by an orchestration server.

[0524] Flowchart 4900B includes an optional operation 4910B for configuring an application comprising a set of distributed nodes to run on an orchestrated system. Flowchart 4900B includes an operation 4920B for running a first module on a first node, the first module having a first output. Flowchart 4900B includes an operation 4930B for running a second module on a second node, the second module using the first output as input. Flowchart 4900B includes an operation 4940B for providing a second output from the second module to a third module running on a third node.

[0525] Flowchart 4900B includes an operation 4950B for determining a replacement node for redeploying the second module by coordinating between the first node and the third node in response to detecting a failure of the second node. In an example, determining the replacement node includes identifying a redundant node preconfigured to receive the first output and operate the second module. The redundant node can be disconnected from any node (e.g., blocked from providing output to any node) until the redundant node is used as a replacement node, such as receiving input and calculating output to maintain the state of the second module, but not connected to any other node. In an example, parameter and status information about the second module can be sent from the second node, the first node, or the third node to the redundant node, such as periodically, whenever an output is generated, etc. In another example, in response to a failure of the redundant node, the second redundant node can be identified to become a replacement node (e.g., for a critical module).

[0526] In an example, determining the redundant node includes determining a set of nodes connected to the second node. The set of nodes may include one or more input nodes or one or more output nodes, such as with a directional indicator. For example, the replacement node may be connected to the first node to receive the output from the first module and connected to the third node to provide the output from the second module to the third module.

[0527] Further operations may include, when generating the first output, saving a redundant state of the second module at the first node. In an example, the orchestration server may initially generate a configuration of the module on the node (e.g., the first module on the first node, etc.). In this example, the orchestration server may be disconnected, for example, before any failure, such as a failure of the second node. Without assistance from the orchestration server, the first node and the third node may coordinate to determine a replacement node. In an example, the second node may be implanted on a virtual machine. The second module may then be instantiated on the replacement node based on the image of the second node on the virtual machine.

[0528] IoT devices and networks

[0529] The above-described techniques can be implemented in conjunction with various device deployments, including in any number of IoT networks and topologies. Therefore, it will be understood that various embodiments of the present technology can involve the coordination of edge devices, fog devices, and middleware, as well as cloud entities, across heterogeneous and homogeneous networks. Some example topologies and arrangements of such networks are provided in the following paragraphs.

[0530] Figure 50 An example domain topology is shown for various Internet of Things (IoT) networks coupled to corresponding gateways via links. The Internet of Things (IoT) is a concept in which a large number of computing devices are interconnected to each other and to the Internet to provide functionality and data collection at a very low level. Therefore, as used herein, IoT devices may include semi-autonomous devices that perform functions (such as sensing or control, etc.) and communicate with other IoT devices and a wider network (such as the Internet).

[0531] IoT devices are physical objects that can communicate over a network and may include sensors, actuators, and other input / output components, such as to collect data from the real-world environment or perform actions. For example, IoT devices may include low-power devices embedded in or attached to everyday objects (such as buildings, vehicles, packages, etc.) to provide an additional level of artificial sensory perception to these objects. Recently, IoT devices have become increasingly popular, and as a result, the applications using these devices have proliferated.

[0532] IoT devices are often limited in terms of memory, size, or functionality, allowing for the deployment of a larger number of devices to achieve similar costs as a smaller number of larger devices. However, IoT devices can be smartphones, laptops, tablets, PCs, or other larger devices. Furthermore, IoT devices can be virtual devices, such as applications on smartphones or other computing devices. IoT devices can include IoT gateways, which are used to couple IoT devices to other IoT devices and to cloud applications for data storage, process control, and so on.

[0533] A network of IoT devices may include commercial and home automation devices such as water systems, power distribution systems, plumbing control systems, factory control systems, light switches, thermostats, locks, cameras, alarms, motion sensors, etc. IoT devices may be accessible through remote computers, servers, and other systems, for example, to control the systems or access data.

[0534] The future growth of the Internet and similar networks may involve a very large number of IoT devices. Accordingly, in the context of the technologies discussed in this article, a large number of innovations for such future networking will address the need for all these layers to grow without obstacles, discover and make accessible connected resources, and support the ability to hide and separate connected resources. Any number of network protocols and communication standards can be used, each of which is designed to address specific goals. In addition, the protocol is part of the structure that supports human-accessible services operating regardless of location, time or space. Innovations include: service delivery and associated infrastructure, such as hardware and software; security enhancements; and service provision based on quality of service (QoS) terms specified in service levels and service delivery agreements. As will be understood, IoT devices and networks using IoT devices and networks such as those introduced in the system example discussed above present a large number of new challenges in heterogeneous connectivity networks that include a combination of wired and wireless technologies.

[0535] Figure 50Specifically, a simplified diagram of a domain topology that can be used for a large number of Internet of Things (IoT) networks is provided, wherein the large number of IoT networks include IoT devices 5004, wherein IoT networks 5056, 5058, 5060, and 5062 are coupled to corresponding gateways 5054 via backbone links 5002. For example, a large number of IoT devices 5004 can communicate with gateway 5054 and can communicate with each other through gateway 5054. To simplify the diagram, not every IoT device 5004 or communication link (e.g., links 5016, 5022, 5028, or 5032) is labeled. Backbone links 5002 can include any number of wired or wireless technologies (including optical networks) and can be part of a local area network (LAN), a wide area network (WAN), or the Internet. In addition, such communication links facilitate optical signal paths between IoT devices 5004 and gateway 5054, including the use of multiplexing / demultiplexing components that facilitate the interconnection of various devices.

[0536] The network topology may include any of a variety of types of IoT networks, such as a mesh network provided using a Bluetooth Low Energy (BLE) link 5022 with network 5056. Other types of IoT networks that may exist include: 5028 wireless local area network (WLAN) network 5058 for communicating with IoT device 5004; cellular network 5060 for communicating with IoT device 5004 via LTE / LTE-A (4G) or 5G cellular network; and low power wide area (LPWA) network 5062, for example, an LPWA network compatible with the LoRaWan specification promulgated by the LoRa Alliance; or IPv6 on a low power wide area network (LPWAN) network compatible with the specification promulgated by the Internet Engineering Task Force (IETF). Further, each IoT network can communicate with an external network provider (e.g., a layer 2 or layer 3 provider) using any number of communication links, such as LTE cellular links, LPWA links, or links based on the IEEE 802.15.4 standard (such as, Each IoT network may also operate with the use of various network and internet application protocols, such as the Constrained Application Protocol (CoAP). Each IoT network may also be integrated with coordinator devices that provide a chain of links that form a cluster tree of linked devices and networks.

[0537] Each of these IoT networks can provide opportunities for new technical features, such as those described herein. Improved technologies and networks can enable exponential growth in devices and networks, including the use of IoT networks as fog devices or fog systems. As the use of such improved technologies grows, IoT networks can be developed to achieve self-management, functional evolution, and collaboration without direct human intervention. Improved technologies can even enable IoT networks to operate without a centralized, controlled system. Accordingly, the improved technologies described herein can be used to automate and enhance network management and operational functions far beyond current implementations.

[0538] In an example, communications between IoT devices 5004 (such as on backbone link 5002) can be protected by a decentralized system for authentication, authorization, and accounting (AAA). In a decentralized AAA system, distributed payment, credit, audit, authorization, and authentication systems can be implemented across interconnected heterogeneous network infrastructures. This allows systems and networks to move toward autonomous operation. In these types of autonomous operations, machines can even enter into human resource contracts and negotiate partnerships with other machine networks. This can allow for common goals and balanced service delivery for outlined planned service level agreements, and implement solutions that provide metering, measurement, traceability, and traceability. The creation of new supply chain structures and methods can enable a large number of services to be generated, valued, and collapsed without any human involvement.

[0539] Such IoT networks can be further enhanced by integrating sensing technologies (such as sound, light, electronic traffic, facial and pattern recognition, smell, and vibration) into the autonomous organization between IoT devices. The integration of sensing systems can allow for systematic and autonomous communication and coordination of service delivery for contracted service objectives, orchestration and quality of service (QoS)-based grouping, and resource convergence. Some of the individual examples of network-based resource processing include the following.

[0540] Mesh network 5056 can be enhanced, for example, by systems that perform inline data-to-information transformations. For example, self-forming chains of processing resources, including multi-link networks, can efficiently distribute the transformation of raw data into information, as well as the ability to differentiate between assets and resources and the associated management of each. Furthermore, trust and service indexes based on appropriate components of infrastructure and resources can be inserted to improve data integrity, quality, and assurance, and deliver metrics of data confidence.

[0541] The WLAN network 5058 may, for example, use a system that performs standard conversion to provide multi-standard connectivity, thereby enabling IoT devices 5004 to communicate using different protocols. Further systems may provide seamless interconnectivity across a multi-standard infrastructure that includes visible Internet resources and hidden Internet resources.

[0542] Communications in the cellular network 5060 may be enhanced, for example, by systems that offload data, systems that extend communications to more remote devices, or both. The LPWA network 5062 may include systems that perform non-Internet Protocol (IP) to IP interconnection, addressing, and routing. Further, each of the IoT devices 5004 may include an appropriate transceiver for wide area communications with that device. Further, each IoT device 5004 may include other transceivers for communicating using additional protocols and frequencies. Figure 52 and Figure 53 This is further discussed in the Communication Environment and Hardware of IoT Processing Devices depicted in .

[0543] Ultimately, clusters of IoT devices can be equipped to communicate with other IoT devices and with cloud networks. This can allow IoT devices to form ad-hoc networks between multiple devices, allowing them to act as a single device, which can be referred to as a fog device. Figure 51 Let's discuss this configuration.

[0544] Figure 51 A cloud computing network is shown communicating with a mesh network of IoT devices (devices 5102) operating as fog devices at the edge of the cloud computing network. The mesh network of IoT devices may be referred to as fog 5120 operating at the edge of cloud 5100. To simplify the diagram, not every IoT device 5102 is labeled.

[0545] The fog 5120 can be considered as a massively interconnected network where several IoT devices 5102 communicate with each other, for example, via radio links 5122. As an example, the interconnected network can use the Open Connectivity Foundation TM The standard is facilitated by the interconnection specification published by the Open Mobile Communications Foundation (OCF). This standard allows devices to discover each other and establish communication for interconnection. Other interconnection protocols may also be used, including, for example, the Optimal Link State Routing (OLSR) protocol, the Better Way to Mobile Ad Hoc Networking (BATMAN) routing protocol, or the OMA Lightweight M2M (LWM2M) protocol, etc.

[0546] Although three types of IoT devices 5102 are shown in this example: gateway 5104, data aggregator 5126, and sensor 5128, any combination of IoT devices 5102 and functions may be used. Gateway 5104 may be an edge device that provides communication between cloud 5100 and fog 5120, and may also provide back-end processing functions for data obtained from sensor 5128 (such as motion data, flow data, temperature data, etc.). Data aggregator 5126 may collect data from any number of sensors 5128 and perform processing functions for analysis. Results, raw data, or both may be passed to cloud 5100 via gateway 5104. Sensor 5128 may be, for example, a complete IoT device 5102 capable of both collecting and processing data. In some cases, sensor 5128 may be more limited in functionality, such as collecting data and allowing data aggregator 5126 or gateway 5104 to process the data.

[0547] Communications from any IoT device 5102 can be passed along a convenient path (e.g., the most convenient path) between any of the IoT devices 5102 to reach the gateway 5104. In these networks, the number of interconnections provides a large amount of redundancy, which allows communications to be maintained even in the event of the loss of several IoT devices 5102. In addition, the use of mesh networks can allow the use of IoT devices 5102 that are very low-power or located at a distance from the infrastructure, as the range of connection to another IoT device 5102 may be much smaller than the range of connection to the gateway 5104.

[0548] The fog 5120 provided from these IoT devices 5102 can be presented to devices in the cloud 5100 (such as the server 5106) as a single device, e.g., a fog device, located at the edge of the cloud 5100. In this example, alerts from the fog device can be sent without being identified as coming from a specific IoT device 5102 within the fog 5120. In this way, the fog 5120 can be viewed as a distributed platform that provides computing and storage resources to perform processing or data-intensive tasks (such as data analytics, data aggregation, and machine learning, etc.).

[0549] In some examples, IoT devices 5102 can be configured using an imperative programming style, for example, with each IoT device 5102 having specific functionality and communication partners. However, IoT devices 5102 forming fog devices can be configured using a declarative programming style, allowing IoT devices 5102 to reconfigure their operation and communication, such as determining required resources in response to conditions, queries, and device failures. As an example, a query from a user at server 5106 regarding the operation of a subset of equipment monitored by IoT devices 5102 can cause fog 5120 devices to select the IoT devices 5102 required to answer the query, such as specific sensors 5128. Data from these sensors 5128 can then be aggregated and analyzed by any combination of sensors 5128, data aggregator 5126, or gateway 5104 before being sent by fog 5120 devices to server 5106 to answer the query. In this example, IoT devices 5102 in fog 5120 can select sensors 5128 to use based on the query, such as adding data from a flow sensor or temperature sensor. Furthermore, if some of the IoT devices 5102 are inoperable, other IoT devices 5102 in the fog 5120 devices may provide similar data if available.

[0550] In an example, various aspects of workload orchestration and operation may be adapted to Figure 51 The various network topologies and approaches depicted. For example, the system can collaborate with IoT device 5102 to establish various workloads that execute in cloud 5100. These workloads can be orchestrated from the edge (e.g., from IoT device 5102) in cloud 5100 or fog 5120, or they can be orchestrated on the edge through cloud 5100 or fog 5120. Such concepts can also be applied to gateway 5104 and data aggregator 5126, as well as other devices and nodes within the network topology.

[0551] In other examples, the operations and functions described above with reference to the systems described above may be embodied by an IoT device machine in the example form of an electronic processing system, within which a set or sequence of instructions may be executed to cause the electronic processing system to perform any of the methods discussed herein according to the examples. The machine may be an IoT device or IoT gateway, including machines embodied by multiple aspects of the following: a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a mobile phone or smartphone, or any machine capable of executing instructions (sequentially or otherwise) specifying actions to be taken by the machine. In addition, although only a single machine may be depicted and referenced in the above examples, such machines should also be considered to include any collection of machines that execute a set (or multiple sets) of instructions, individually or in combination, to perform any one or more of the methods discussed herein. In addition, these examples and similar examples of processor-based systems should be considered to include any collection of one or more machines controlled or operated by a processor (e.g., a computer) to execute instructions, individually or in combination, to perform any one or more of the methods discussed herein.

[0552] Figure 52 A diagram illustrates a cloud computing network or cloud 5200 in communication with several Internet of Things (IoT) devices. Cloud 5200 may represent the Internet, or it may be a local area network (LAN) or a wide area network (WAN), such as a company's private network. IoT devices may include any number of different types of devices grouped in various combinations. For example, traffic control group 5206 may include IoT devices along a city street. These IoT devices may include stop lights, traffic flow monitors, cameras, weather sensors, and so on. Traffic control group 5206 or other subgroups may communicate with cloud 5200 via wired or wireless links 5208 (such as LPWA links, optical links, and so on). Furthermore, wired or wireless subnets 5212 may allow IoT devices to communicate with each other, such as via a local area network, a wireless local area network, and so on. IoT devices may use another device (such as gateway 5210 or 5228) to communicate with a remote location (such as cloud 5200); IoT devices may also use one or more servers 5230 to facilitate communication with cloud 5200 or gateway 5210. For example, one or more servers 5230 can act as intermediate network nodes to support local edge cloud or fog implementations between local area networks. Furthermore, the depicted gateway 5228 can operate in a cloud-gateway-many edge device configuration, such as with various IoT devices 5214, 5220, 5224 that are constrained to, or dynamic with respect to, the allocation and use of resources in the cloud 5200.

[0553] Other example groups of IoT devices may include a remote weather station 5214, a local information terminal 5216, an alarm system 5218, an automated teller machine 5220, an alarm panel 5222, or a mobile vehicle (such as an emergency vehicle 5224 or other vehicle 5226), etc. Each of these IoT devices may be connected to other IoT devices, to a server 5204, to another IoT fog device or system (not shown but Figure 51 ), or communicate with a combination thereof. Groups of these IoT devices can be deployed in a variety of residential, commercial, and industrial settings (including both private and public environments).

[0554] As from Figure 52 As can be seen, a large number of IoT devices can communicate through the cloud 5200. This allows different IoT devices to autonomously request information or provide information to other devices. For example, a group of IoT devices (e.g., traffic control group 5206) can request the current weather forecast from a group 5214 of remote weather stations that can provide forecasts without human intervention. In addition, an emergency vehicle 5224 can be warned by an automated teller machine 5220 that a theft is in progress. When the emergency vehicle 5224 is traveling toward the automated teller machine 5220, it can access the traffic control group 5206 to request that the location be cleared, for example, by illuminating a red light for a sufficient time to block cross traffic flow at the intersection so that the emergency vehicle 5224 can enter the intersection unimpeded.

[0555] Clusters of IoT devices (such as remote weather stations 5214 or traffic control groups 5206) may be equipped to communicate with other IoT devices and with the cloud 5200. This may allow IoT devices to form an ad hoc network between multiple devices, allowing them to act as a single device, which may be referred to as a fog device or system (e.g., as described above with reference to Figure 51 described).

[0556] Figure 53 is a block diagram of examples of components that may be present in an IoT device 5350 for implementing the techniques described herein. The IoT device 5350 may include any combination of components shown in the examples or referenced in the disclosure above. These components may be implemented as ICs, portions of ICs, discrete electronic devices, or other modules, logic, hardware, software, firmware, or combinations thereof suitable for use in the IoT device 5350, or as components otherwise incorporated within the chassis of a larger system. Additionally, Figure 53 The block diagram is intended to depict a high-level view of the components of the IoT device 5350. However, some of the components shown may be omitted, additional components may be present, and different arrangements of the components shown may occur in other implementations.

[0557] IoT device 5350 may include a processor 5352, which may be a microprocessor, a multi-core processor, a multi-threaded processor, an ultra-low voltage processor, an embedded processor, or other known processing element. Processor 5352 may be part of a system on a chip (SoC), in which processor 5352 and other components are formed into a single integrated circuit or a single package, such as the Edison processor from Intel. TM (Edison TM ) or Galileo TM (Galileo TM )SoC board. As an example, the processor 5352 may include a Core Architecture TM (Core TM ) processors (such as Quark TM 、Atom TM , i3, i5, i7, or MCU class processors), or available from However, any number of other processors may be used, such as processors available from Advanced Micro Devices, Inc. (AMD) of Sunnyvale, California, MIPS-based designs from MIPS Technologies, Inc. of Sunnyvale, California, ARM-based designs licensed from ARM Holdings, Inc., or processors obtained from customers, licensees, or adopters of the aforementioned companies. The processor may include units such as: The company's A5-A10 processors, from Snapdragon Technologies TM (Snapdragon TM ) processor or OMAP from Texas Instruments TM processor.

[0558] The processor 5352 can communicate with the system memory 5354 via an interconnect 5356 (e.g., a bus). Any number of memory devices can be used to provide a fixed amount of system memory. As an example, the memory can be a random access memory (RAM) designed according to the Joint Electron Devices Engineering Council (JEDEC), such as DDR or mobile DDR standards (e.g., LPDDR, LPDDR2, LPDDR3, or LPDDR4). In various implementations, the individual memory devices can be any number of different package types, such as a single die package (SDP), a dual die package (DDP), or a quad die package (Q17P). In some examples, these devices can be directly soldered to the motherboard to provide a simpler solution, while in other examples, the devices are configured as one or more memory modules, which in turn are coupled to the motherboard via a given connector. Any number of other memory implementations can be used, such as other types of memory modules, for example, different types of dual in-line memory modules (DIMMs), including but not limited to microDIMMs or MiniDIMMs.

[0559] To provide persistent storage of information (such as data, applications, operating systems, etc.), storage 5358 may also be coupled to processor 5352 via interconnect 5356. In an example, storage 5358 may be implemented via a solid-state disk drive (SSDD). Other devices that can be used for storage 5358 include flash memory cards (such as SD cards, microSD cards, xD picture cards, etc.) and USB flash drives. In low-power implementations, storage 5358 may be on-die memory or registers associated with processor 5352. However, in some examples, storage 5358 may be implemented using a micro hard disk drive (HDD). Furthermore, in addition to or in lieu of the described technologies, any number of new technologies may be used for storage 5358, such as resistive memory, phase change memory, holographic memory, or chemical memory, etc.

[0560] Components can communicate via interconnect 5356. Interconnect 5356 can include any number of technologies, including Industry Standard Architecture (ISA), Extended ISA (EISA), Peripheral Component Interconnect (PCI), Peripheral Component Interconnect Extended (PCIx), PCI Express (PCIe), or any number of other technologies. Interconnect 5356 can be a proprietary bus, such as used in SoC-based systems. Other bus systems can be included, such as an I2C interface, an SPI interface, a point-to-point interface, a power bus, and the like.

[0561] The interconnect 5356 can couple the processor 5352 to a mesh transceiver 5362 for communication with other mesh devices 5364, for example. The mesh transceiver 5362 can use any number of frequencies and protocols, such as 2.4 gigahertz (GHz) transmission under the IEEE 802.15.4 standard, using a method such as that described by Special Interest Group Definition Low Energy (BLE) standard, or standards, etc. Any number of radios configured for a particular wireless communication protocol may be used for connection to the mesh device 5364. For example, a WLAN unit may be used to implement Wi-Fi according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard. TM Additionally, wireless wide area communications, such as according to cellular or other wireless wide area protocols, may occur via the WWAN unit.

[0562] Mesh transceiver 5362 can communicate using a variety of standards or radios for communication at different ranges. For example, IoT device 5350 can communicate with nearby devices (e.g., within about 10 meters) using a local transceiver based on BLE or another low-power radio to save power. Mesh devices 5364 that are farther away (e.g., within about 50 meters) can be contacted using ZigBee or other medium-power radios. The two communication technologies can occur over a single radio at different power levels, or can occur through separate transceivers, such as a local transceiver using BLE and a separate mesh transceiver using ZigBee.

[0563] A wireless network transceiver 5366 may be included to communicate with devices or services in the cloud 5300 via a local area network protocol or a wide area network protocol. The wireless network transceiver 5366 may be an LPWA transceiver that complies with the IEEE 802.15.4 or IEEE 802.15.4g standards, etc. The IoT device 5350 may use LoRaWAN developed by Semtech and the LoRa Alliance. TM (Long Range Wide Area Network) to communicate over a wide area. The techniques described herein are not limited to these technologies and can be used with any number of other cloud transceivers that enable long-range, low-bandwidth communications, such as Sigfox and other technologies. In addition, other communication techniques can be used, such as time-division channel hopping as described in the IEEE 802.15.4e specification.

[0564] In addition to the systems mentioned for the mesh transceiver 5362 and the wireless network transceiver 5366 as described herein, any number of other radio communications and protocols may be used. For example, the radio transceivers 5362 and 5366 may include LTE or other cellular transceivers that use spread spectrum (SPA / SAS) communications to achieve high-speed communications. Additionally, any number of other protocols may be used, such as LTE for medium-speed communications and provisioning network communications. network.

[0565] Radio transceivers 5362 and 5366 may include radios compatible with any number of 3GPP (3rd Generation Partnership Project) specifications, particularly Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), and Long Term Evolution-Advanced Plus (LTE-A Pro). It may be noted that radios compatible with any number of other fixed, mobile, or satellite communication technologies and standards may be selected. These may include, for example, any cellular wide area radio communication technology, which may include, for example, a 5th Generation (5G) communication system, a Global System for Mobile Communications (GSM) radio communication system, a General Packet Radio Service (GPRS) radio communication technology, or an Enhanced Data Rate for GSM Evolution (EDGE) radio communication technology, a UMTS (Universal Mobile Telecommunications System) communication technology, and in addition to the standards listed above, any number of satellite uplink technologies may be used for the wireless network transceiver 5366, including, for example, radios compliant with standards published by the ITU (International Telecommunication Union) or ETSI (European Telecommunications Standards Institute). The examples provided herein may therefore be understood to be applicable to various other communication technologies, both existing and yet to be developed.

[0566] A network interface controller (NIC) 5368 may be included to provide wired communication to the cloud 5300 or to other devices, such as mesh devices 5364. The wired communication may provide an Ethernet connection, or may be based on other types of networks, such as a controller area network (CAN), a local interconnect network (LIN), a device network (DeviceNet), a control network (ControlNet), a data highway+, a field bus (PROFIBUS), or an industrial Ethernet (PROFINET), etc. An additional NIC 5368 may be included to allow connection to a second network, for example, a NIC 5368 providing communication to the cloud via Ethernet, and a second NIC 5368 providing communication to other devices via another type of network.

[0567] Interconnect 5356 can couple processor 5352 to external interface 5370, which is used to connect external devices or subsystems. External devices may include sensors 5372, such as accelerometers, level sensors, flow sensors, optical light sensors, camera sensors, temperature sensors, global positioning system (GPS) sensors, pressure sensors, barometric pressure sensors, etc. External interface 5370 can further be used to connect IoT device 5350 to actuators 5374 (such as power switches, valve actuators, audible sound generators, visual warning devices, etc.).

[0568] In some optional examples, various input / output (I / O) devices may be present within the IoT device 5350 or may be connected to the IoT device 5350. For example, a display or other output device 5384 may be included to display information, such as sensor readings or actuator positions. Input devices 5386 (such as a touch screen or keys) may be included to accept input. The output device 5384 may include any number of audio or visual display forms, including: simple visual output, such as a binary status indicator (e.g., an LED); multi-character visual output; or more complex output, such as a display screen (e.g., an LCD screen) with output of characters, graphics, multimedia objects, etc. generated or produced from the operation of the IoT device 5350.

[0569] The battery 5376 can power the IoT device 5350, but in an example where the IoT device 5350 is installed in a fixed location, the IoT device 5350 can have a power source coupled to the power grid. The battery 5376 can be a lithium-ion battery, a metal-air battery (such as a zinc-air battery, an aluminum-air battery, a lithium-air battery), or the like.

[0570] A battery monitor / charger 5378 may be included in the IoT device 5350 to track the state of charge (SoCh) of the battery 5376. The battery monitor / charger 5378 may be used to monitor other parameters of the battery 5376 to provide failure prediction, such as the state of health (SoH) and state of function (SoF) of the battery 5376. The battery monitor / charger 5378 may include a battery monitoring integrated circuit, such as the LTC4020 or LTC2990 from Linear Technologies, the ADT7488A from ON Semiconductor in Phoenix, Arizona, or the UCD90xxx family of ICs from Texas Instruments in Dallas, Texas. The battery monitor / charger 5378 may communicate information on the battery 5376 to the processor 5352 via the interconnect 5356. The battery monitor / charger 5378 may also include an analog-to-digital (ADC) converter that allows the processor 5352 to directly monitor the voltage of the battery 5376 or the current flowing from the battery 5376. Battery parameters may be used to determine actions that the IoT device 5350 may perform, such as transmission frequency, mesh network operation, sensing frequency, and the like.

[0571] A power block 5380 or other power source coupled to the grid can be coupled to a battery monitor / charger 5378 to charge the battery 5376. In some examples, the power block 5380 can be replaced with a wireless power receiver to obtain power wirelessly, for example, via a loop antenna in the IoT device 5350. Wireless battery charging circuitry (such as the LTC4020 chip from Linear Technology Corporation of Milpitas, California, etc.) can be included in the battery monitor / charger 5378. The specific charging circuit selected depends on the size of the battery 5376 and, therefore, the required current. Charging can be performed using the Airfuel standard promulgated by the Airfuel Alliance, the Qi wireless charging standard promulgated by the Wireless Power Consortium, or the Rezence charging standard promulgated by the Alliance for Wireless Power, among others.

[0572] Storage 5358 may include instructions 5382 in the form of software, firmware, or hardware commands for implementing the techniques described herein. While such instructions 5382 are illustrated as code blocks included in memory 5354 and storage 5358, it will be appreciated that any of the code blocks may be replaced with hardwired circuitry, such as built into an application specific integrated circuit (ASIC).

[0573] In an example, the instructions 5382 provided via the memory 5354, storage 5358, or processor 5352 may be embodied as a non-transitory machine-readable medium 5360 that includes code for instructing the processor 5352 to perform electronic operations in the IoT device 5350. The processor 5352 may access the non-transitory machine-readable medium 5360 via the interconnect 5356. For example, the non-transitory machine-readable medium 5360 may be provided by a processor 5352 for executing an instruction to execute ... Figure 53 The non-transitory machine-readable medium 5360 may include instructions for instructing the processor 5352 to perform, for example, a specific sequence of actions or flow of actions as described with reference to the flowchart(s) and block diagram(s) of the operations and functions depicted above.

[0574] In a further example, a machine-readable medium also includes any tangible medium that is capable of storing, encoding, or carrying instructions for execution by a machine and causing the machine to perform any one or more of the methods of the present disclosure, or that is capable of storing, encoding, or carrying data structures utilized by or associated with such instructions. "Machine-readable media" may therefore include, but is not limited to, solid-state memories, optical media, and magnetic media. Specific examples of machine-readable media include non-volatile memories, including, by way of example, but not limited to: semiconductor memory devices (e.g., electrically programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The instructions embodied by the machine-readable medium may be transmitted or received further over a communications network using a transmission medium, via a network interface device, using any of several transmission protocols (e.g., HTTP).

[0575] It should be understood that the functional units or capabilities described in this specification may be referred to or labeled as components or modules, thereby particularly emphasizing the independence of their implementation. Such components may be embodied in any number of software or hardware forms. For example, a component or module may be implemented as a hardware circuit comprising: a customized very large scale integration (VLSI) circuit or gate array, an off-the-shelf semiconductor such as a logic chip, a transistor, or other discrete components. A component or module may also be implemented in a programmable hardware device (such as a field programmable gate array, programmable array logic, a programmable logic device, etc.). A component or module can also be implemented in software for execution by various types of processors. The identified component or module of the executable code may, for example, include one or more physical or logical blocks of computer instructions, which may, for example, be organized into objects, processes, or functions. However, the executable files of the identified components or modules do not have to be physically together, but may include different instructions stored in different locations, which, when logically combined together, comprise the component or module and achieve the claimed purpose for the component or module.

[0576] In practice, a component or module of executable code can be a single instruction or many instructions, and can even be distributed over several different code segments, between different programs, and across several memory devices or processing systems. Specifically, some aspects of the described processes (such as code rewriting and code analysis) may occur on a processing system (e.g., a computer in a data center) that is different from the processing system in which the code is deployed (e.g., a computer embedded in a sensor or robot). Similarly, operational data may be identified and illustrated herein within a component or module and can be embodied in any suitable form and organized within any suitable type of data structure. Operational data may be collected as a single data set, or may be distributed over different locations (including over different storage devices), and may exist, at least in part, solely as electronic signals on a system or network. Components or modules may be passive or active, including agents for performing the desired functions.

[0577] Example

[0578] Additional examples of embodiments of the presently described methods, systems, and devices include the following non-limiting configurations. Each of the following non-limiting examples can stand alone or be combined with one or more of the other examples provided below or throughout the present disclosure in any permutation or combination.

[0579] In Example 1, a method for operating a software-defined industrial system includes: establishing a corresponding functional definition of the software-defined industrial system, wherein the software-defined industrial system is used to interface with multiple devices, wherein the multiple devices include corresponding sensors and corresponding actuators; and using the corresponding functional definition to operate the software-defined industrial system.

[0580] In Example 2, the subject matter of Example 1 includes: establishing a dynamic data model to define properties of a plurality of components of the software-defined industrial system; and updating the dynamic data model based on operational metadata associated with the plurality of components.

[0581] In Example 3, the subject matter of Example 2 includes: wherein the plurality of components include respective applications, devices, sensors, or architecture definitions.

[0582] In Example 4, the subject matter of Examples 2-3 includes: wherein the plurality of components comprises a device, wherein the device represents a population of sensors.

[0583] In Example 5, the subject matter of Examples 2-4 includes: wherein the dynamic data model is updated to indicate changes to the dynamic data model in a subset of components in the plurality of components, and wherein the dynamic data model is updated based on a change in resource availability or an error condition occurring in the subset of components.

[0584] In Example 6, the subject matter of Examples 2-5 includes: wherein establishing the dynamic data model includes defining mandatory fields and restrictions on changes to the dynamic data model.

[0585] In Example 7, the subject matter of Examples 2-6 includes: wherein the operational metadata represents a probabilistic estimate of a value associated with a component of the plurality of components.

[0586] In Example 8, the subject matter of Examples 2-7 includes: querying a component among the plurality of components for metadata expansion rules; receiving a response from the component in response to the query; wherein updating the dynamic data model is also based on the metadata expansion rules and a confidence or relevance score associated with an update response data field.

[0587] In Example 9, the subject matter of Examples 2-8 includes: monitoring the data flows from the plurality of components to identify the operational metadata; detecting one or more patterns from the plurality of components; identifying changes to the dynamic data model based on the detected one or more patterns; wherein updating the dynamic data model includes incorporating the identified changes.

[0588] In Example 10, the subject matter of Examples 2-9 includes performing system operations in an edge, fog, or cloud network based on an updated dynamic data model.

[0589] In Example 11, the subject matter of Examples 1-10 includes: defining at least one condition in the software-defined industrial system for data model evaluation; acquiring data from multiple sensors in the software-defined industrial system; identifying at least one pattern, rule, or threshold for data model modification; evaluating data from the multiple sensors using at least one identified pattern, rule, or identified threshold; defining modifications to the data model based on the at least one identified pattern, rule, or identified threshold; and incorporating modifications to the data model for the multiple sensors and data streams associated with the multiple sensors.

[0590] In Example 12, the subject matter of Example 11 includes: requesting approval of the data model modification from a data model administrator; and obtaining approval of the data model modification from the data model administrator; wherein, in response to receiving the approval of the data model modification, incorporating the modification into the data model.

[0591] In Example 13, the subject matter of Examples 11-12 includes implementing changes to data processing operations in the software-defined industrial system based on the data model modification.

[0592] In Example 14, the subject matter of Examples 1-13 includes establishing an extended set of orchestrator logic rules for function blocks executed across a distributed resource pool in the software-defined industrial system.

[0593] In Example 15, the subject matter of Example 14 includes performing dynamic discovery of network bandwidth, resource capacity, current status, and control application constraints for the distributed resource pool.

[0594] In Example 16, the subject matter of Examples 14-15 includes establishing an orchestration with the corresponding legacy device via a shim that interfaces with the corresponding legacy application.

[0595] In Example 17, the subject matter of Examples 14-16 includes, wherein the extended orchestrator rule set includes one or more of: application cycle time, application runtime, application input / output signal dependency, or application process order.

[0596] In Example 18, the subject matter of Examples 14-17 includes: evaluating functional block application timing dependencies of the application deployment based on the application cycle, runtime dependencies of the application deployment, and the current state of the application deployment; and allocating corresponding applications of the application deployment among nodes of the software-defined industrial system based on the evaluated functional block application timing dependencies.

[0597] In Example 19, the subject matter of Examples 14-18 includes: monitoring various functional blocks of an application deployment; updating optimizations and predictive forecasts based on current and historical data; and orchestrating the execution of one or more of the various functional blocks in a distributed resource pool according to a control policy in response to detecting a system anomaly in one or more of the various functional blocks.

[0598] In Example 20, the subject matter of Example 19 includes determining whether the control strategy is feasible, wherein, in response to determining that the control strategy is feasible, performing orchestrated execution of one or more of the respective functional blocks.

[0599] In Example 21, the subject matter of Example 20 includes, in response to determining that the control strategy is not feasible, implementing a degraded or disconnected control strategy for at least a portion of the one or more of the respective functional blocks.

[0600] In Example 22, the subject matter of Examples 14-21 includes: wherein the distributed resource pool includes applications across one or more of: a single application running on a single native device, where a second redundant application is available on a nearby native device; multiple coordinated applications running in multiple native devices; multiple coordinated applications running on a single virtual machine, where the virtual machine runs on a single embedded device or server; multiple coordinated applications running across multiple virtual machines, where each virtual machine runs in a dedicated embedded device or server; multiple coordinated applications across multiple containers contained in a virtual machine, where the virtual machine runs in a dedicated embedded device or server; or multiple coordinated applications across multiple containers, where the containers run on multiple embedded devices or servers.

[0601] In Example 23, the subject matter of Examples 14-22 includes: wherein establishing an extended set of orchestrator logic rules for function blocks executed across distributed resource pools includes: identifying application-specific dependencies; dynamically creating orchestration groups of distributed and dependent applications based on the identified dependencies; predicting orchestration events; detecting the predicted orchestration events; and optimizing resource placement in response to detecting the predicted orchestration events.

[0602] In Example 24, the subject matter of Example 23 includes: wherein predicting the orchestration event comprises dynamically analyzing and simulating network bandwidth in an example scenario and analyzing occurrence of the orchestration event in the example scenario.

[0603] In Example 25, the subject matter of Examples 1-24 includes: establishing communication with a legacy component, wherein the legacy component is a legacy software module or a legacy hardware device; and establishing communication with an orchestration component, wherein the orchestration component is an orchestration software module or an orchestration hardware device; and establishing an organized orchestration to control and distribute workload between the orchestration component and the legacy component.

[0604] In Example 26, the subject matter of Example 25 includes: establishing an orchestration shim to configure a legacy software module, and wherein the orchestration shim is adapted to provide a custom configuration to the legacy software module; and communicating directly with the legacy software module based on the custom configuration for the control and distribution of the workload.

[0605] In Example 27, the subject matter of Example 26 includes: transmitting the custom configuration to the legacy software module through an application programming interface (API) of the orchestration shim; and transmitting legacy module communication information from the legacy software module through the API of the orchestration shim, wherein the communication with the legacy software module is further performed using the legacy module communication information.

[0606] In Example 28, the subject matter of Examples 26-27 includes: transmitting a second configuration to the programmable software module via an application programming interface (API) of the orchestration software module; and directly communicating with the programmable software module based on the second configuration for the control and distribution of the workload.

[0607] In Example 29, the subject matter of Examples 25-28 includes: establishing, by the orchestrated hardware device, an organized orchestration with the legacy hardware device based on telemetry collected from an agent of the orchestrated hardware device, the telemetry indicating available resources of the legacy hardware device; and deploying, by the agent of the orchestrated hardware device, a workload based on the organized orchestration to the legacy hardware device.

[0608] In Example 30, the subject matter of Examples 25-29 includes: establishing an organized orchestration with an orchestrated hardware device based on telemetry collected from an agent of the orchestrated hardware device, the telemetry indicating available resources of the orchestrated hardware device; and deploying a workload based on the organized orchestration to the orchestrated hardware device via the agent of the orchestrated hardware device.

[0609] In Example 31, the subject matter of Examples 25-30 includes: receiving, at an orchestration engine of an orchestrator, a description of available resources from respective orchestratable devices, wherein the description of available resources is based on telemetry received from the respective orchestratable devices; organizing a hierarchy of orchestrations defined from the orchestrator to the respective orchestratable devices based on the description of available resources; and distributing workloads from the orchestration engine to the respective orchestratable devices based on the hierarchy of orchestrations.

[0610] In Example 32, the subject matter of Example 31 includes: wherein the hierarchical structure is a functional hierarchy of orchestration, wherein the hierarchical structure defines application orchestration by using sub-orchestration software modules, and wherein the sub-orchestration software modules include respective software modules for network orchestration, virtual machine orchestration, task orchestration, and storage orchestration.

[0611] In Example 33, the subject matter of Examples 31-32 includes: wherein the orchestrated hierarchy is a single-level hierarchy, and wherein the orchestration engine assigns a subset of the respective orchestratable devices to run a portion of a corresponding workload.

[0612] In Example 34, the subject matter of Examples 31-33 includes: wherein the hierarchical structure of the orchestration is a multi-level hierarchical structure, the multi-level hierarchical structure including a sub-orchestration having a corresponding orchestration engine on a middle level of the multi-level hierarchical structure, wherein the orchestration engine and the orchestrator operate at a top level of the multi-level hierarchical structure, and wherein the corresponding orchestration engine operates to coordinate the collection of the telemetry and the distribution of the workload among the respective orchestratable devices at a bottom level of the multi-level hierarchical structure.

[0613] In Example 35, the subject matter of Example 34 includes: wherein the groups of orchestratable devices are organized into clusters, and wherein the sub-orchestrators coordinate the collection of the telemetry and the distribution of the workload among the clusters.

[0614] In Example 36, the subject matter of Examples 31-35 includes: wherein the orchestrated hierarchical structure is a multi-level hierarchical structure, the multi-level hierarchical structure including a master orchestratable device at a middle level of the multi-level hierarchical structure, and slave nodes at a bottom level of the multi-level hierarchical structure, wherein the orchestration engine and the orchestrator run at a top level of the multi-level hierarchical structure, and wherein the master orchestratable device includes a corresponding agent for coordinating the collection of the telemetry and the distribution of the workload among the slave nodes.

[0615] In Example 37, the subject matter of Example 36 includes: wherein each cluster is organized based on pairing each master orchestration device with at least one slave node; wherein the corresponding agent coordinates distribution of the workload among the each cluster.

[0616] In Example 38, the subject matter of Examples 36-37 includes performing detection, discovery, and deployment of individual slave nodes at the bottom level of a multi-level hierarchy.

[0617] In Example 39, the subject matter of Examples 31-38 includes: collecting software data, hardware data, and network data from components of an organized orchestration, the organized orchestration comprising the traditional components and the orchestratable components; performing monitoring, by an orchestration server, based on the collected software data, hardware data, and network data; and providing feedback and control from the orchestration server to the components of the organized orchestration to control the organized orchestration in response to the monitoring.

[0618] In Example 40, the subject matter of Examples 1-39 includes defining and deploying a self-describing control application and software modules for the software-defined industrial system, wherein the self-describing control application includes a plurality of self-describing programmable software modules.

[0619] In Example 41, the subject matter of Example 40 includes: creating a module manifest for describing characteristics of the programmable software module; defining an application specification based on the definition and connection of available features in the programmable software module; defining options and alternatives for the operation of the programmable software module; and selecting the programmable software module based on the options and alternatives.

[0620] In Example 42, the subject matter of Example 41 includes simulating and evaluating operation of the programmable software module in a simulated application setting, wherein the selection based on the programmable software module is based on a result of the simulated application setting.

[0621] In Example 43, the subject matter of Example 42 includes: wherein the operations for simulating and evaluating the operation of the programmable software module include: determining available application and software module configurations using an application specification and one or more module manifests; defining a plurality of orchestration scenarios through a feature controller; executing the application module and at least one alternative application module with (a plurality of) defined options using a simulator to implement the plurality of orchestration scenarios; evaluating execution results of the application module and the at least one alternative application module based on hardware performance and user input; and generating corresponding scores for the execution results of the application module and the at least one alternative application module.

[0622] In Example 44, the subject matter of Examples 42-43 includes, wherein scenarios associated with the execution results are automatically incorporated for use in the application based on respective scores.

[0623] In Example 45, the subject matter of Example 1 includes: receiving data from a field device (e.g., a sensor), such as at an IO converter; converting the data from the field device according to a field device bus protocol; sending the converted data to a field device abstraction bus; receiving a control signal from a control device via the field device abstraction bus; and sending an electrical signal to the field device based on the control signal.

[0624] In Example 46, the subject matter of Example 1 includes: receiving data from multiple field devices (e.g., sensors) via multiple corresponding IO converters, such as at a sensor bus; sending the data to one or more control functions; receiving one or more control signals from the one or more control functions based on the data, and sending the one or more control signals to each of the multiple IO converters.

[0625] In Example 47, the subject matter of Example 46 includes receiving information from an IO converter mode controller, and facilitating assigning the IO converter to the field device based on the information received from the IO converter mode controller.

[0626] In Example 48, the subject matter of Example 1 includes: saving information about multiple alarms of an industrial control system; analyzing data, context, or alarm configuration of the multiple alarms from the information; determining alarm flow similarity from the information; detecting alarm events at two or more alarms; preventing the two or more alarms from being issued; and generating clustered alarms for the two or more alarms that were prevented from being issued.

[0627] In Example 49, the subject matter of Example 48 includes suggesting changes to one or more of the plurality of alarms or suggesting a new alarm.

[0628] In Example 50, the subject matter of Example 1 includes managing autonomous creation of a new closed-loop workload algorithm.

[0629] In Example 51, the subject matter of Example 50 includes performing a quality or sensitivity assessment of the new algorithm relative to a current process (eg, an industrial control system process).

[0630] In Example 52, the subject matter of Examples 50-51 includes autonomously establishing operational constraint boundaries.

[0631] In Example 53, the subject matter of Examples 50-52 includes autonomously evaluating the safety of the new algorithm relative to existing processes.

[0632] In Example 54, the themes of Examples 50-53 include: the value of autonomously assessing the broader process.

[0633] In Example 55, the subject matter of Examples 50-54 includes autonomously assessing deployment feasibility of the system in a control environment.

[0634] In Example 56, the subject matter of Examples 50-55 includes physically deploying or monitoring the new application control policy.

[0635] In Example 57, the subject matter of Examples 50-56 includes integrating the new algorithm into a lifecycle management system.

[0636] In Example 58, the subject matter of Examples 50-57 includes integrating the new algorithm into the scrapping process.

[0637] In Example 59, the subject matter of Examples 50-58 includes performing Examples 51-58 in order.

[0638] In Example 60, the subject matter of Example 1 includes: determining computing requirements of edge control nodes in an industrial control system (e.g., a ring deployment), such as at an orchestration server: receiving an indication to activate a CPU of one or more edge control nodes; and sending an authenticated activation code to the edge control node having the CPU to be activated.

[0639] In Example 61, the subject matter of Example 1 includes receiving an authenticated activation code at an edge control node; authenticating the code at the edge control node and activating a CPU of the edge control node using a microprocessor (MCU) (eg, a low-performance processor).

[0640] In Example 62, the subject matter of Examples 60-61 includes performing Examples 60-61 at a ring deployment of edge control nodes arranged by an orchestration system of an industrial control system.

[0641] In Example 63, the subject matter of Example 61 includes receiving, at the edge control node, an update from the orchestration server to deactivate a CPU or place the CPU in a low power state.

[0642] Example 64 is at least one machine-readable medium comprising instructions that, when executed by a computing system, cause the computing system to perform any of Examples 1-63.

[0643] Example 65 is an apparatus comprising corresponding apparatus for performing any of Examples 1-63.

[0644] Example 66 is a software-defined industrial system, comprising a corresponding device and corresponding circuits in the corresponding device, wherein the corresponding circuits are configured to perform the operations of any one of Examples 1-63.

[0645] Example 67 is an apparatus comprising circuitry configured to perform the operations of any of Examples 1-63.

[0646] In Example 68, the subject matter of Example 67 includes: wherein the apparatus is a gateway that enables connection to an adapted plurality of field devices, other device networks, or other network deployments.

[0647] In Example 69, the subject matter of Examples 67-68 includes: wherein the apparatus is a device operably coupled to at least one sensor and at least one actuator.

[0648] In Example 70, the subject matter of Examples 67-69 includes: wherein the apparatus is an edge control node device adapted to connect to a plurality of field devices.

[0649] In Example 71, the subject matter of Examples 67-70 includes: wherein the apparatus is an intelligent I / O controller device adapted to connect to a plurality of field devices.

[0650] In Example 72, the subject matter of Examples 67-71 includes wherein the apparatus is a basic I / O controller device adapted to connect to a plurality of field devices.

[0651] In Example 73, the subject matter of Examples 67-72 includes: wherein the device is a control server computing system adapted to connect to a plurality of network systems.

[0652] In Example 74, the subject matter of Examples 67-73 includes: wherein the apparatus is a control processing node computing system adapted to connect to a plurality of networked systems.

[0653] Example 75 is a network system comprising respective devices connected within a fog or cloud network topology, the respective devices comprising circuitry configured to perform the operations of any of Examples 1-63.

[0654] In Example 76, the subject matter of Example 75 includes: wherein the respective devices are connected via a real-time service bus.

[0655] In Example 77, the subject matter of Examples 75-76 includes: wherein the network topology includes controller, storage, and compute functionality for the software-defined industrial system via redundant host pairs.

[0656] In Example 78, the subject matter of Examples 75-77 includes: wherein the network topology includes controller, storage, and compute functionality for the software-defined industrial system via separate physical hosts.

[0657] Example 79 is an edge control node of an industrial control system, comprising: an input / output (IO) subsystem for receiving signals from field devices and generating IO data; and a system on a chip comprising: a networking component communicatively coupled to a network; a microcontroller (MCU) for converting the IO data from the IO subsystem and sending the converted data to an orchestration server via the network through the networking component; and a central processing unit (CPU) that is initially in an inactive state and is configured to become active in response to receiving an activation signal from the orchestration server through the networking component at the edge control node.

[0658] In Example 80, the subject matter of Example 79 includes: wherein the activation state of the CPU includes a low power mode and a high power mode.

[0659] In Example 81, the subject matter of Examples 79-80 includes wherein the CPU is further configured to receive a deactivation signal from the orchestration server after a period of time in the active state, and in response, return to the inactive state.

[0660] In Example 82, the subject matter of Examples 79-81 includes: wherein the edge control node is one of a plurality of edge control nodes in the industrial control system, the plurality of edge control nodes including at least one edge control node having an inactive CPU after the CPU is activated.

[0661] In Example 83, the subject matter of Examples 79-82 includes, wherein the CPU is activated based on the orchestration server determining that the CPU is to be activated to satisfy a control policy for the industrial control system.

[0662] In Example 84, the subject matter of Examples 79-83 includes, wherein the networking component is a time-sensitive network Ethernet switch.

[0663] In Example 85, the subject matter of Examples 79-84 includes: wherein the network has a ring topology having a bridge device connecting the network to the orchestration server.

[0664] In Example 86, the subject matter of Examples 79-85 includes: wherein the activation signal is received at the CPU directly from the MCU.

[0665] In Example 87, the subject matter of Examples 79-86 includes: wherein the CPU is further configured to receive processing instructions from the orchestration server, the CPU being configured to execute the processing instructions while in the activated state.

[0666] Example 88 is at least one non-transitory machine-readable medium comprising instructions that, when executed by a processor of an orchestration server, cause the processor to: receive input / output (10) data, the 10 data being received via a bridge connecting the orchestration server to an edge control node, wherein the 10 data is converted at a microcontroller (MCU) of the edge control node from data generated at an 10 subsystem to data packets sent by a networking component; send an authenticated activation code to the edge control node to activate a central processing unit (CPU) of the edge control node, wherein the CPU is initially placed in an unactivated state; and send processing instructions to the CPU for execution.

[0667] In Example 89, the subject matter of Example 88 includes: wherein the operation further causes the processor to determine a computing requirement of an edge control node in an industrial control system including the edge control node, and wherein the CPU is activated based on the orchestration server determining that activating the CPU satisfies a control policy for the industrial control system.

[0668] In Example 90, the subject matter of Examples 88-89 includes, wherein the operations further cause the processor to receive an instruction to activate the CPU of the edge control node in the industrial control system.

[0669] In Example 91, the subject matter of Examples 88-90 includes: wherein, before activating the CPU, the authenticated activation code is authenticated by the MCU

[0670] In Example 92, the subject matter of Examples 88-91 includes, wherein the operations further cause the processor to send a deactivation code from the orchestration server to the CPU to deactivate the CPU.

[0671] In Example 93, the subject matter of Examples 88-92 includes: wherein the edge control node has a ring topology network having a bridge device connecting the network to the orchestration server.

[0672] Example 94 is an industrial control system comprising: a ring network including a plurality of edge control nodes; an orchestration server; a bridge connecting the orchestration server to the ring network; and wherein the plurality of edge control nodes include, a first edge control node comprising: a system on a chip including: a microcontroller (MCU) for converting input / output (IO) data from an IO subsystem and sending the converted data to the orchestration server through the bridge via a networking component; and a processor in an initially inactive state for: receiving an activation signal from the orchestration server; and changing to an active state in response to receiving the activation signal.

[0673] In Example 95, the subject matter of Example 94 includes wherein the processor is further configured to receive a deactivation signal from the orchestration server after a period of the activated state, and in response, return to the inactive state.

[0674] In Example 96, the subject matter of Examples 94-95 includes, wherein the processor is activated based on an orchestration server determining that the processor satisfies a control policy for the industrial control system.

[0675] In Example 97, the subject matter of Examples 94-96 includes: wherein the activation signal is received at the processor directly from the MCU.

[0676] In Example 98, the subject matter of Examples 94-97 includes, wherein the plurality of edge control nodes includes a second edge control node, wherein the second processor remains inactive after the processor of the first edge control node is activated.

[0677] In Example 99, the subject matter of Examples 94-98 includes, wherein the orchestration server is further configured to send the processing instructions to the processor for execution.

[0678] In Example 100, the subject matter of Examples 94-99 includes: wherein the processor is a central processing unit (CPU).

[0679] Example 101 is at least one machine-readable medium comprising instructions that, when executed by a processing circuit, cause the processing circuit to perform operations to implement any of Examples 79-100.

[0680] Example 102 is an apparatus comprising means for implementing any one of Examples 79-100.

[0681] Example 103 is a system for implementing any of Examples 79-100.

[0682] Example 104 is a method for implementing any of Examples 79-100.

[0683] Example 105 is a device comprising processing circuitry adapted to: identify operational aspects of available software modules, the available software modules being adapted to perform functional operations in a control system environment; identify operational characteristics from a module inventory, wherein the operational characteristics define an environment for the available software modules to execute a control system application; select a software module from the available software modules based on the identified operational aspects of the available software modules and the operational characteristics identified from the module inventory; and cause execution of the selected software module in the control system environment, wherein the execution occurs according to an application specification of the control system application.

[0684] In example 106, the subject matter of example 105 includes: wherein the operational aspects of the available software modules relate to one or more of: a communication interface, a starting parameter, a platform requirement, a dependency, a deployment requirement, or a signature.

[0685] In Example 107, the subject matter of Examples 105-106 includes: the processing circuit is further adapted to: generate the application specification for the control system application based on the operating characteristics and the selected software module; wherein the application specification defines a value of a control parameter of the selected software module.

[0686] In Example 108, the subject matter of Example 107 includes wherein the application specification indicates a connection from the selected software module to a second selected software module.

[0687] In Example 109, the subject matter of Examples 105-108 includes the processing circuit being further adapted to: evaluate execution of the selected software module in the control system environment using at least two different hardware architectures; and perform efficiency measurements on operations performed using the at least two different hardware architectures.

[0688] In Example 110, the subject matter of Examples 105-109 includes: wherein the control system application and each software module are displayed as a visual representation in a graphical user interface, wherein the visual representation is used to establish a relationship of one or more inputs or outputs of the software module within the control system application, wherein the inputs or outputs of the software module include using one or more of the following: sensors, actuators, or controllers.

[0689] In Example 111, the subject matter of Examples 105-110 includes: wherein the device is an orchestration device, wherein the orchestration device is operably coupled to a plurality of execution devices that execute software modules in the control system environment, and wherein execution of a selected software module via at least one execution device affects functional operation of one or more control devices in the control system environment.

[0690] In Example 112, the subject matter of Example 111 includes wherein the processing circuit is further adapted to coordinate execution of the selected software modules using an orchestrated control strategy in the control system environment.

[0691] In Example 113, the subject matter of Examples 105-112 includes: wherein the processing circuit is further adapted to: select a plurality of software modules, the plurality of software modules comprising the selection of the software module; and connect the plurality of software modules to each other according to the operating characteristic.

[0692] Example 114 is a method comprising: identifying operational aspects of available software modules, the available software modules being suitable for performing functional operations in a control system environment; identifying operational characteristics from a module inventory, wherein the operational characteristics define an environment for the available software modules to execute a control system application; selecting a software module from the available software modules based on the identified operational aspects of the available software modules and the operational characteristics identified from the module inventory; and causing execution of the selected software module in the control system environment, wherein the execution occurs according to an application specification of the control system application.

[0693] In example 115, the subject matter of example 114 includes: wherein the operational aspects of the available software modules relate to one or more of: a communication interface, a starting parameter, a platform requirement, a dependency, a deployment requirement, or a signature.

[0694] In Example 116, the subject matter of Examples 114-115 includes generating the application specification for the control system application based on the operating characteristics and the selected software module; wherein the application specification defines a value of a control parameter of the selected software module, and wherein the application specification indicates a connection from the selected software module to a second selected software module.

[0695] In Example 117, the subject matter of Examples 114-116 includes evaluating execution of the selected software module in a control system environment using at least two different hardware architectures; and identifying efficiency measures for operations performed using the at least two different hardware architectures.

[0696] In Example 118, the subject matter of Examples 114-117 includes: wherein the control system application and each software module are displayed as a visual representation in a graphical user interface, wherein the visual representation is used to establish a relationship of one or more inputs or outputs of the software module within the control system application, wherein the inputs or outputs of the software module include using one or more of the following: sensors, actuators, or controllers.

[0697] In Example 119, the subject matter of Examples 114-118 includes: wherein the method is performed by an orchestration device, wherein the orchestration device is operably coupled to a plurality of execution devices that execute software modules in the control system environment, and wherein execution of the selected software module via at least one execution device affects functional operation of one or more control devices in the control system environment.

[0698] In Example 120, the subject matter of Example 119 includes utilizing an orchestrated control policy in the control system environment to coordinate execution of the selected software modules.

[0699] In Example 121, the subject matter of Examples 119-120 includes selecting a plurality of software modules for use in the control system environment, the plurality of software modules comprising a selection of the software modules; and interconnecting the plurality of software modules based on the operational characteristic.

[0700] Example 122 is at least one non-transitory machine-readable storage medium comprising instructions that, when executed by processing circuitry of a device, cause the processing circuitry to perform a plurality of operations, the plurality of operations comprising: identifying operational aspects of available software modules, the available software modules being suitable for performing functional operations in a control system environment; identifying operational characteristics from a module inventory, wherein the operational characteristics define an environment for the available software modules to execute a control system application; selecting a software module from the available software modules based on the identified operational aspects of the available software modules and the operational characteristics identified from the module inventory; and causing execution of the selected software module in the control system environment, wherein the execution occurs according to an application specification of the control system application.

[0701] In Example 123, the subject matter of Example 122 includes: wherein the operational aspects of the available software modules relate to one or more of: a communication interface, a starting parameter, a platform requirement, a dependency, a deployment requirement, or a signature.

[0702] In Example 124, the subject matter of Examples 122-123 includes that the operation further includes: generating the application specification for the control system application based on the operating characteristics and the selected software module; wherein the application specification defines the value of the control parameter of the selected software module, and wherein the application specification indicates a connection from the selected software module to a second selected software module.

[0703] In Example 125, the subject matter of Examples 122-124 includes, the operations further comprising: evaluating execution of the selected software module in the control system environment using at least two different hardware architectures; and performing efficiency measurements on the operations performed using the at least two different hardware architectures.

[0704] In Example 126, the subject matter of Examples 122-125 includes: wherein the control system application and each software module are displayed as a visual representation in a graphical user interface, wherein the visual representation is used to establish a relationship of one or more inputs or outputs of the software module within the control system application, wherein the inputs or outputs of the software module include using one or more of the following: sensors, actuators, or controllers.

[0705] In Example 127, the subject matter of Examples 122-126 includes: wherein the plurality of operations are performed by an orchestration device, wherein the orchestration device is operably coupled to a plurality of execution devices that execute software modules in the control system environment, and wherein execution of the selected software module via at least one execution device affects functional operation of one or more control devices in the control system environment.

[0706] In Example 128, the subject matter of Example 127 includes, the operations further comprising: utilizing an orchestration control policy in the control system environment to coordinate execution of the selected software modules.

[0707] In Example 129, the subject matter of Examples 127-128 includes that the operation further includes: selecting a plurality of software modules for use in the control system environment, the plurality of software modules comprising the selection of the software modules; and connecting the plurality of software modules to each other according to the operating characteristics.

[0708] Example 130 is at least one machine-readable medium comprising instructions that, when executed by a processing circuit, cause the processing ci...

Claims

1. An edge control node of an industrial control system, comprising: An input / output (IO) subsystem for receiving signals from field devices and generating IO data; as well as System on chip, comprising: a networking component communicatively coupled to a network; a microcontroller (MCU) configured to convert the IO data from the IO subsystem and send the converted data to an orchestration server via a network through the networking component; and A central processing unit (CPU), initially in an inactive state, is configured to become active in response to receiving an activation signal from the orchestration server at the edge control node through the networking component.

2. The edge control node according to claim 1, wherein: The activation state of the CPU includes a low power mode and a high power mode.

3. The edge control node according to claim 1, wherein: The CPU is further configured to receive a deactivation signal from the orchestration server after a period of time in the active state, and in response, return to the inactive state.

4. The edge control node according to claim 1, wherein: The edge control node is one of a plurality of edge control nodes in the industrial control system, the plurality of edge control nodes including at least one edge control node having an inactive CPU after the CPU is activated.

5. The edge control node according to claim 1, wherein: The CPU is activated based on the orchestration server determining that the CPU is to be activated to satisfy a control policy for the industrial control system.

6. The edge control node according to claim 1, wherein: The networking component is a time-sensitive network Ethernet switch.

7. The edge control node according to claim 1, wherein: The network has a ring topology with a bridge device connecting the network to the orchestration server.

8. The edge control node according to claim 1, wherein: The activation signal is received at the CPU directly from the MCU.

9. The edge control node according to claim 1, wherein: The CPU is further configured to receive a processing instruction from the orchestration server, and the CPU is configured to execute the processing instruction when in the activated state.

10. A method performed by an orchestration server in an industrial control system, comprising: receiving input / output (IO) data, the IO data being received via a bridge connecting the orchestration server to an edge control node, wherein the IO data is converted at a microcontroller (MCU) of the edge control node from data generated at an IO subsystem into data packets sent by a networking component; sending an authenticated activation code to the edge control node to activate a central processing unit (CPU) of the edge control node, wherein the CPU is initially placed in an inactive state; and Processing instructions are sent to the CPU for execution.

11. The method of claim 10, further comprising: Computational requirements of an edge control node in an industrial control system including the edge control node are determined, and wherein the CPU is activated based on the orchestration server determining that activating the CPU satisfies a control policy for the industrial control system.

12. The method of claim 10, further comprising: An instruction is received to activate the CPU of the edge control node in the industrial control system.

13. The method according to claim 10, wherein The authenticated activation code is authenticated by the MCU before the CPU is activated.

14. The method of claim 10, further comprising: A deactivation code is sent from the orchestration server to the CPU to deactivate the CPU.

15. The method according to claim 10, wherein The edge control node is a node of a ring topology network having a bridge device connecting the network to the orchestration server.

16. At least one machine-readable medium comprising instructions for operation of a computing system, said instructions, when executed by a machine, causing said machine to perform the method of any one of claims 10-15.

17. An industrial control system comprising: a ring network comprising a plurality of edge control nodes; Orchestration server; a bridge connecting the orchestration server to the ring network; as well as The plurality of edge control nodes include a first edge control node, including: System on chip, comprising: a microcontroller MCU, configured to convert input / output IO data from an input / output IO subsystem and send the converted data to the orchestration server via the bridge through a networking component; and The processor, which is initially inactive, is used to: receiving an activation signal from the orchestration server; and The state is changed to an active state in response to receiving the activation signal.

18. The industrial control system according to claim 17, wherein: The processor is further configured to receive a deactivation signal from the orchestration server after a period of time in the active state, and in response, return to the inactive state.

19. The industrial control system according to claim 17, wherein: The processor is activated based on the orchestration server determining that activating the processor satisfies a control policy for the industrial control system.

20. The industrial control system according to claim 17, wherein: The activation signal is received at the processor directly from the MCU.

21. The industrial control system according to claim 17, wherein: The plurality of edge control nodes include a second edge control node, wherein a second processor remains inactive after the processor of the first edge control node is activated.

22. The industrial control system according to claim 17, wherein: The orchestration server is further configured to send processing instructions to the processor for execution.

23. The industrial control system according to claim 17, wherein: The processor is a central processing unit CPU.

Citation Information

Patent Citations

  • Electronic marketing system and method

    AU2005242171A1

  • Product activation / registration and offer eligibility

    CN101911039A