Symbiotic predictive digital twins with machine learning automation: integrated methods for continuous monitoring and model management
Patent Information
- Application Number
- PCT/US2025/025613
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2025-04-21
- Publication Date
- 2025-12-04
AI Technical Summary
Existing Digital Twin Platforms face challenges in cost-effectiveness, operationalization scalability, and practicality in implementing Level 3 Digital Twins due to complexities in integrating advanced predictive analytics with data collection infrastructures, leading to high costs and extended time-to-value.
A symbiotic integration with established data collection infrastructure using Timeseries Data Historian and Contextualization Framework Technology, enabling modular design and automation for efficient machine learning model selection and management, streamlining Level 3 Digital Twin operationalization across diverse industries.
Reduces costs and accelerates time-to-value by integrating predictive analytics into digital twin design from the outset, ensuring value-driven insights and efficient deployment of predictive models with continuous monitoring and anomaly detection.
Smart Images

Figure US2025025613_04122025_PF_FP_ABST
Abstract
Description
SYMBIOTIC PREDICTIVE DIGITAL TWINS WITH MACHINE LEARNING AUTOMATION: INTEGRATED METHODS FOR CONTINUOUS MONITORING AND MODEL MANAGEMENTCROSS-REFERENCE TO RELATED APPLICATION
[0001] This Application is an International Application filed under the Patent Cooperation Treaty that claims the priority / benefit of U.S. Provisional Application No. 63 / 636,454 filed April 19, 2024, the disclosure of which is expressly incorporated by reference herein in its entirety.BACKGROUND1. Field of the Disclosure
[0002] This technology relates to methods and systems for designing and implementing predictive analytics-enabled digital twins (Level 3), which are achieved through a symbiotic integration with established Data Collection Infrastructure. This infrastructure includes components such as Timeseries Data Historian and Contextualization Framework Technology.2. Background Information
[0003] Digital Twins are virtual replicas of physical systems that can be used for simulation, analysis, and optimization. They have been widely used in various industries, such as manufacturing, transportation, healthcare, and utilities to improve efficiency, reduce costs, and increase safety.
[0004] Level 3 Digital Twins utilize advanced predictive analytics, surpassing conventional data analysis to enhance operational efficiency and minimize downtime. The implementation of Level 3 Digital Twins on a large industrial scale presents challenges attributed to substantial costs and an extended time-to-value.
[0005] While modern Data Collection Infrastructure systems provide a framework for collecting, organizing, and augmenting raw data through abstract concepts such as elements, element templates, attributes, and event frames, they lack native support for advanced predictive analytics, such as Machine Learning or Artificial Intelligence. Although industrial companies have made significant investments in Data CollectionInfrastructure, they face difficulties in implementing Level 3 Digital Twins due to the complexities associated with using such abstract concepts, creating appropriate digital structures, and integrating advanced analytics at scale. This highlights the need for additional technology layers capable of designing, deploying and operationalizing Level 3 Digital Twins.
[0006] Currently available Digital Twin Platforms face limitations in terms of cost-effectiveness and operationalization scalability. Note that operationalization scalability refers not only to the ability to ingest, analyze, and process large volume of events, but also to scalability in terms of implementation practicality, complexity, cost, and duration.Existing Digital Twin Platforms treat customer's data collection infrastructures as merely a data source into their platform or ignoring it altogether, requiring additionalinvestment in data collection, streaming and digitization frameworks, which can impede the adoption of their technology by many industrial companies.
[0007] Finally, when it comes to operationalizing and maintaining many predictive analytics models, these platforms lack practicality, making it difficult to achieve a balance between nimbleness, complexity, and automation.SUMMARY
[0008] The present disclosure proposes a paradigm shift in digitization process by symbiotically integrating with pre-established data collection infrastructure, making it a core element. Leveraging modular design, automation and progressive complexity approach enables efficient machine learning model selection and management. Unlike traditional methods, predictive analytics are integrated into digital twin design from the outset, ensuring value-driven insights throughout. Automation spans ETL processes, predictive monitoring strategy discovery, model selection, training, and deployment, simplifying complex data science tasks into intuitive configuration items. This adaptable approach caters to diverse industries and equipment, offering a powerful blend of nimbleness, complexity, and automation for predictive analytics. By streamlining Level 3 Digital Twin operationalization, the present disclosure reduces costs and accelerates time-to-value.
[0009] The present disclosure introduces a multi-layer computing system designed for deployment in both on-premises and cloud environments, leveraging an established Data Collection Infrastructure (DCI) comprising a Timeseries Data Historian (TDH) and Contextualization Framework Technology (CFT). This infrastructure ensures the historization of sensor data from industrial equipment into the TDH, facilitated by hardware and software components such as sensors, PLC, firewalls, switches, routers, and network interfaces. The CFT enhances raw tags from the Data Historian Server using abstract concepts like elements, element templates, attributes, and event frames, supporting data ingestion from various sources and facilitating event processing based on user-defined rules. It also provides an API for native read / write operations, reflecting standard practices within modern DCI technology across targeted industries.
[0010] Moreover, the disclosure introduces a conceptual layer tailored to resonate with diverse industries and equipment, which is translated into requisite constructs within the CFT through disclosed software layers and methods. This approach streamlines the operationalization and management of Statistical and Machine Learning (ML) models, bridging the conceptual and technological realms to enhance accessibility and efficiency in deploying predictive analytics solutions across various industrial contexts.
[0011] Central to the disclosure is the concept of the "Asset", representing operating equipment or industrial processes undergoing digitization, comprised of "Asset Components" and "Predictive Modules". These entities facilitate the design of digital twins mirroring physical assets and employ state-of-the-art algorithms such as multivariate regression and pattern recognition for predictive monitoring. The disclosure also introduces the concept of "Asset Shape", offering a template of templates for digital structure and automating the generation and deployment of Level 3 Digital Twins based on this shape, providing an efficient and vendor-agnostic solution.
[0012] The implementation of these concepts is realized through a series of software layers, including Repository, Internal Storage, Domain, Service, Web API, Web Client, and ML Execution Engine. These layers facilitate seamless integration, efficient data management, and real-time monitoring of deployed models. The architecture emphasizes interoperability and modularity, allowing for flexibility in integrating with diverse DCI providers and adapting to evolving business needs or technological advancements.
[0013] The present disclosure introduces several features to enhance visibility and monitoring of deployed predictive models across the asset fleet. The Model Coverage dashboard provides metrics on model coverage, gaps, and complexity distribution. Outliers within the fleet for SQC and Regression models are automatically identified, ensuring accurate monitoring of all assets.
[0014] The ML Model Management Repository extracts configuration details of predictive modules, while the Data ETL Repository retrieves the latest sensor readings for targets and predictors. The ML Execution Engine processes this information to perform predictions and writes them back to designated tags in the Timeseries Data Historian.
[0015] Continuous monitoring involves triggering events based on predefined rules within the CFT, generating events for SQC and Regression models if targets fall outside control limits or prediction intervals. Anomalies are flagged based on these events. Moreover, users can take corrective actions or remedial measures to address the flagged / detected anomalies, e.g., replacing sensors or parts, changing process parameters, etc.
[0016] The Active Anomalies dashboard displays near real-time active anomalous outcomes, calculating the Rate of Change for anomalous targets. Users can drill down into asset summaries to access detailed information about predictive module anomaly states and sensor readings overlaid on a 3D model of the asset.
[0017] The Advanced Data Exploration tool streamlines anomaly investigation by providing data exploration tools and insights into anomalous behavior. Users can access this tool through the Predictive Models dashboard, conducting bivariate analysis and peer fleet comparisons.
[0018] The tool also enables retraining of statistical and machine learning models on the fly, ensuring up-to-date and optimized models for continuous monitoring.
[0019] Embodiments are directed to a method of producing predictive digital twin of an industrial machine having at least one physical component. The method includes creating an asset template, which is a digital representation of the at least one physical component, from a base asset template stored in at least one memory; creating an asset component template from a base asset component template stored in the at least one memory; creating a predictive module template, which includes predictors and targets, from a base predictive module template stored in the at least one memory; and creating an asset shape, which defines a digital structure for an asset represented in the asset template, by selecting the asset component template and the predictive module template.
[0020] In further embodiments, the method can include executing the predictive digital twin to identify anomalies in the industrial machine. Further, the executing of the predictive digital twin can include executing a machine learning process.
[0021] In accordance with embodiments, the at least one physical component can include data collection infrastructure, which can include at least one of a sensor, a programmable logic controller, a firewall, a switch, a router and a network interface.
[0022] Embodiments are directed to a predictive digital twin system that includes at least one memory device; a computer processing device; a data collection infrastructure associated with at least one physical component to collect operational data and to transmit the collected operational data to the at least one memory device; a plurality of templates accessible from the at least one memory for configuring assets, asset components and predictive modules; and an asset shape, which defines a digital structure of the at least one physical component that is represented in an asset template through selected asset component templates and predictive module templates.
[0023] According to embodiments, the data collection infrastructure can include at least one of a sensor, a programmable logic controller, a firewall, a switch, a router and a network interface.
[0024] In accordance with other embodiments, the data collection infrastructure may include a plurality of sensors, which are arranged on the at least one physical component, and the collected operational data may include at least one of temperature, pressure or speed.
[0025] In embodiments, the predictive digital twin can be configured as a model of an industrial machine comprising the at least one physical component.
[0026] According to other embodiments, the predictive digital twin may be executable to identify anomalies in the industrial machine. Further, in identifying anomalies in the industrial machine, a machine learning process can be executed.
[0027] Embodiments are directed to a predictive digital twin system that includes a processor; and at least one memory device storing a set of computer readable instructions, which, when executed by the processor, cause the system to: create an asset template, which is a digital representation of at least one physical component, from a base asset template stored in at least one memory; create an asset component template from a base asset component template stored in the at least one memory; create a predictive module template, which includes predictors and targets, from a base predictive module template stored in the at least one memory; and create an asset shape, which defines a digital structure for an asset represented in the asset template, by selecting the asset component template and the predictive module template.
[0028] According to further embodiments, the stored set of computer readable instructions, which, when executed by the processor, can cause the system to execute the predictive digital twin to identify anomalies in the industrial machine. Moreover, the stored set of computer readable instructions, which, when executed by the processor, may further cause the system to execute a machine learning process.
[0029] In accordance with still yet other embodiments, the at least one physical component may include data collection infrastructure, which can include at least one of a sensor, a programmable logic controller, a firewall, a switch, a router and a network interface.BRIEF DESCRIPTION OF THE DRAWINGS
[0030] FIG. 1 depicts a schematic outlining the logical sequence of steps involved in implementing the present disclosure, along with the bounded contexts associated with each step.
[0031] FIG. 2 presents a schematic illustrating the technology layers and data flow specifically within the Digital Templates Bounded Context.
[0032] FIG. 3 illustrates the classes and their corresponding attributes for the domain entities within the Digital Templates Bounded Context.
[0033] FIG. 4 depicts the exemplary user interface alongside pertinent technology layers involved in the process of Creating or Editing an Asset Template.
[0034] FIG. 5 depicts the exemplary user interface alongside pertinent technology layers involved in the process of Creating or Editing an Asset Component Template.
[0035] FIG. 6 depicts the exemplary user interface alongside pertinent technology layers involved in the process of Creating or Editing a Predictive Module Template.
[0036] FIG. 7 depicts the exemplary user interface alongside pertinent technology layers involved in the process of Creating or Editing an Asset Shape.
[0037] FIG. 8 presents a schematic illustrating the technology layers and data flow specifically within the Asset Information Repository Bounded Context.
[0038] FIG. 9 illustrates the classes and their corresponding attributes for the domain entities within the Asset Information Repository Bounded Context.
[0039] FIG. 10A and 10B depict the exemplary user interface alongside pertinent technology layers involved in the process of Creating or Editing an Asset.
[0040] FIG. 11A, 11B and 11C depict the exemplary user interface alongside pertinent technology layers involved in the process of Mapping Tags for an Asset.
[0041] FIG. 12 illustrates the exemplary user interface representing a fully implemented Asset upon completion of Step S3.
[0042] FIG. 13 presents a schematic illustrating the technology layers and data flow specifically within the Machine Learning & Model Management Bounded Context.
[0043] FIG. 14A and 14B illustrate the classes and their corresponding attributes for the domain entities within the Machine Learning & Model Management Bounded Context.
[0044] FIG. 15 depicts the relational database diagram showcasing pertinent tables for machine learning models and their assignments to assets.
[0045] FIG. 16 depicts the exemplary user interface alongside pertinent technology layers involved in the process of Monitoring Strategy Auto-Discovery.
[0046] FIG. 17 illustrates a flow chart diagram elucidating the logic employed for recommending monitoring strategies for all Targets defined within a given Predictive Module Template.
[0047] FIG. 18 depicts the exemplary user interface alongside pertinent technology layers involved in the process of configuring Regression and Pattern Recognition within a Predictive Module Template.
[0048] FIG. 19 illustrates the exemplary user interface along with the relevant technology layers utilized for generating the Model Management Summary Dashboard.
[0049] FIG. 20 depicts the exemplary user interface alongside pertinent technology layers involved in the process of batch auto-training of statistical and machine learning models configured for a given Asset.
[0050] FIG. 21A, 21B, 21C, and 21D illustrate flow chart diagrams elucidating the logic employed for batch auto-training of statistical and machine learning models configured for a given Asset.
[0051] FIG. 22 depicts the exemplary user interface alongside pertinent technology layers involved in the process of deploying auto-trained predictive models for a given Asset.
[0052] FIG. 23 depicts the exemplary user interface alongside pertinent technology layers involved in the process of manually training an SQC model for a given Target.
[0053] FIG. 24A and 24B depict the exemplary user interface alongside pertinent technology layers involved in the process of manually training a Regression model for a given Target.
[0054] FIG. 25A and 25B depict the exemplary user interface alongside pertinent technology layers involved in the process of manually training a Clustering model for a given Predictive Module.
[0055] FIG. 26 presents a schematic illustrating the technology layers and data flow specifically within the ML Execution Engine Bounded Context.
[0056] FIG. 27 presents a schematic illustrating the technology layers and data flow specifically within the Advanced Data Exploration Bounded Context.
[0057] FIG. 28 illustrates the classes and their corresponding attributes for the domain entities within the Advanced Data Exploration Bounded Context.
[0058] FIG. 29 illustrates the exemplary user interface along with the relevant technology layers utilized for generating the Active Anomalies Dashboard.
[0059] FIG. 30 illustrates the exemplary user interface along with the relevant technology layers utilized for generating the Asset Anomaly Summary Dashboard.
[0060] FIG. 31 depicts the exemplary user interface alongside pertinent technology layers involved in the process of advanced data exploration (Step S8) for a given Predictive Module. This exploration occurs through the Predictive Models dashboard and the Temporal Data Analysis functionality.
[0061] FIG. 32A and 32B depict the exemplary user interface alongside pertinent technology layers involved in the process of advanced data exploration (Step S8) for a given Predictive Module. This exploration occurs through the Fleet Comparison functionality.
[0062] FIG. 33A and 33B illustrate the necessary infrastructure setup for the described system, covering both on-premises and cloud environments. This includes an overview of data flow, as well as the hardware, software, and network components involved.DETAILED DESCRIPTION
[0063] The present disclosure introduces a multi-layer computing system suitable for deployment in both on-premises and cloud environments as depicted in FIG. 33 and 33b. In either environment, the system relies on established data collection infrastructure (DCI), incorporating a Timeseries Data Historian (TDH) and Contextualization Framework Technology (CFT). The pre-existing DCI ensures the historization of sensor data from industrial equipment into the TDH. Through hardware and software components such as sensors, PLC, firewalls, switches, routers, and network interfaces, the DCI facilitates a read-only data flow from the Control Network (sensors on industrial equipment) to the Business Network (Data Historian Server).
[0064] Within the Control Network, sensors on industrial machines continuously collect data on key operational parameters like temperature, pressure, or speed. This data is funneled to Programmable Logic Controllers (PLCs), which act as an initial data processing hub, aggregating, and refining data from various sensors into structured information packets. These packets then traverse secure network channels, protected by firewalls placed between the Control and Business Networks. These firewalls are specifically configured to permit only certain preapproved data transmissions, safeguarding against unauthorized access, and ensuring data remains read-only to the recipient. As the data travels across the network, switches and routers play crucial roles in managing and directing traffic. Switches ensure smooth traffic flow within network segments, directing packets swiftly. Routers bridge separate network segments, including the secured links between Control and Business Networks, guiding data packets efficiently towards the Business Network's Data Historian Server. Network interfaces act as gatekeepers at both ends of the data flow. On the Control Network side, they format and send off the PLC-generated data packets. Conversely, on the Business Network side, they receive these packets, delivering them to the Data Historian Server. This server, designed for the storage of time-series data, facilitates effective data querying, reporting, and historical analysis of the industrial process data.
[0065] The CFT offers capabilities for organizing and enhancing raw tags from the Data Historian Server using abstract concepts like elements, element templates, attributes, and event frames. It can ingest data from various sources, including real-time tags, relational databases, and static data, while supporting the augmentation of tags with calculated data (soft sensors) through custom formulas. Additionally, the CFT facilitates the processing and generating of large volume of events based on user-defined rules triggered either on a scheduled basis or by sensor readings, sending notifications via email or other channels. Furthermore, it includes features for templatization of attribute sets and template inheritance concepts. Importantly, the CFT provides an Application Programming Interface (API) exposing native read / write operations.These characteristics of an established DCI serve as prerequisites for the functionality described in this disclosure, reflecting standard practices within modern DCI technology and common usage across targeted industries.
[0066] The disclosed system architecture described in FIG. 33A and 33B includes both on-premises and cloud configurations, centralized around a server-grade computer (server hardware) and a relational database on a Relational Database Server, such as SQL Server. In the on-premises setup (FIG. 33A), the system is integrated within the Business Network. It connects to the existing Data Collection Infrastructure (DCI), facilitating data operations further detailed in the disclosure. Users access the system through a browser using the HTTP protocol, ensuring ease of use and accessibility. The system's Relational Database is linked to the contextualization framework via an OLEDB connection. Conversely, in the cloud-based architecture (FIG. 33B), the disclosed system is hosted in the cloud, leveraging modern infrastructure to provide flexibility and scalability. Connectivity between the system and the Business Network is secured with additional firewalls, ensuring safe data transfer and system access across the network.
[0067] The present disclosure introduces a multi-layer computing system suitable for deployment in both on-premises and cloud environments. In either scenario, the system relies on established data collection infrastructure (DCI), incorporating a Timeseries Data Historian (TDH) and Contextualization Framework Technology (CFT). The pre-existing DCI ensures the historization of sensor data from industrial equipment into the TDH. Through hardware and software components such as sensors, PLC, firewalls, switches, routers, and network interfaces, the DCI facilitates a read-only data flow from the Control Network (sensors on industrial equipment) to the Business Network (Data Historian Server). The CFT offers capabilities for organizing and enhancing raw tags from the Data Historian Server using abstract concepts like elements, element templates, attributes, and event frames. It can ingest data from various sources, including real-time tags, relational databases, and static data, while supporting the augmentation of tags with calculated data (soft sensors) through custom formulas. Additionally, the CFT facilitates the processing and generating of large volume of events based on user defined rules triggered either on a scheduled basis or by sensor readings, sending notifications via email or other channels. Furthermore, it includes features for templatization of attribute sets and template inheritance concepts. Importantly, the CFT provides an Application Programming Interface (API) exposing native read / write operations. These characteristics of an established DCI serve as prerequisites for the functionality described in this disclosure, reflecting standard practices within modern DCI technology and common usage across targeted industries.
[0068] The present disclosure introduces a conceptual layer crafted to resonate with diverse industries and industrial equipment, ensuring intuitive usability for end-users. These concepts are subsequently translated into requisite constructs native to the pre-existing Contextualization Framework Technology (CFT) through disclosed software layers and methods. Moreover, the disclosure streamlines the operationalization and management of Statistical and Machine Learning (ML) models through disclosed methods. By bridging the conceptual and technological realms, this approach enhances accessibility and efficiency in deploying predictive analytics - enabled digital twin solutions across various industrial contexts.
[0069] In the present disclosure, the fundamental concept driving digital twin implementation is the "Asset", serving as the cornerstone representing the overarching class of operating equipment or industrial processes undergoing digitization. Comprised of "Asset Components", these entities facilitate the design of digital twins that closely mirror physical assets and associated real-lifeprocesses in a logical manner. They dynamically define the structure of the Asset, incorporating sensors and static reference data fields. The "Predictive Modules" represent subsystems leveraging Machine Learning to anticipate potential abnormal data behavior, encapsulating all relevant variables, and employing state-of-the-art algorithms such as multivariate regression and pattern recognition. A Predictive Module contains a list of Predictors - independent variables that explain the data behavior for the given module. It also contains a list of Targets - variables that describe the condition of the given subsystem and are targeted for predictive monitoring. A Target can be set up to be monitored in real-time with a combination of univariate statistical method and multi-variate machine learning algorithms. Furthermore, Pattern Recognition can be enabled for any given Predictive Module. This uses advanced clustering algorithms to identify abnormal patterns in the data. Through a modular design and progressive complexity approach, operationalizing predictive analytics models is simplified, featuring automated model selection and novel model management methodologies to achieve an optimal balance of nimbleness, complexity, and automation.
[0070] Introducing the concept of "Asset Shape", the disclosure outlines the anticipated digital structure for any given asset type. This concept offers "Template of Templates" functionality, facilitating components and predictive modules reusability, along with asset 3D scans data layouts. Going beyond conventional approaches, the present invention automates the generation and deployment of Level 3 Digital Twins based on the Asset Shape. This streamlines the digitization process across diverse industries, delivering an efficient and effective equipment vendor-agnostic solution.
[0071] The implementation of the conceptual layer is realized through a series of software layers:1. The Repository Layer serves as a bi-directional interface tailored for the given established DCI Technology. Leveraging accessible APIs, this layer adeptly manages data read / write operations while remaining cognizant of the DCI's inherent constructs. By fostering a symbiotic relationship between the disclosed system and the existing DCI, the Repository Layer facilitates seamless integration and interoperability, ensuring efficient data management and utilization within the system.2. The Internal Storage consists of a relational database utilized for storing trained statistical and machine learning models, along with their corresponding asset assignments. The Repository Layer is responsible for interfacing with this storage, utilizing ORM (Object-Relational Mapping) technology, such as Entity Framework to persist the models and assignments into the relational database efficiently. Moreover, Internal Storage seamlessly integrates with the customer's Contextualization Framework Technology (CFT) Server through native protocols such as OLEDB, facilitating table lookups for CFT's native attributes.3. The Domain Layer represents the disclosed system domain's concepts, rules, and logic. It encapsulates domain entities (classes with associated attributes and methods), embodying the essential business concepts and behaviors. The Domain Layer defines the structure and behavior of the domain model, enforcing business invariants and constraints. It contains the domain logic responsible for executing business operations and ensuring data consistency and integrity. The domain entities are instantiated and data-bound by the Repository Layer from the established CFT server.4. The Service Layer serves as a mediator between the Domain Layer and the Repository Layer, responsible for orchestrating interactions and enforcing business rules. It encapsulates business logic that involves multiple domain entities collaborating to fulfill specific use cases, ensuring consistency and integrity throughout operations. The Service Layer manages transactions, ensuring that operations within a single use case are either all successful or all fail together, and orchestrates the application workflow by coordinating domain operations and handling error scenarios.5. The Web API Layer provides a web service that adheres to the principles of Representational State Transfer (REST). The RESTful API layer serves as an interface between the Web Client and the underlying layers, providing access to resources and operations via standard HTTP methods. It defines the structure and behavior of the API endpoints, each representing a resource or an action within the system. The RESTful API layer handles requests from clients, interpreting HTTP methods like GET, POST, PUT, DELETE to perform corresponding CRUD (Create, Read, Update, Delete) operations on resources. It also manages the JSON serialization and deserialization of data formats to facilitate communication to and from the Web Client. Additionally, the RESTful API layer includes features like authentication and authorization to ensure security.6. The Web Client provides modern, highly intuitive user interfaces that streamline and support all aspects of the digitization process. It collects user inputs; constructs JSON objects and communicates with the Web API Layer through AJAX calls.7. The ML Execution Engine integrates with the established CFT through native APIs and leverages the Service, Domain, and Repository Layers for evaluating machine learning models deployed in near-real time, running on a scheduled basis (i.e., Windows Service). It also writes the predictive model results to the established TDH (Timeseries Data Historian) tags through available native APIs.
[0072] From an implementation perspective, the Internal Storage component is realized as a Relational Database residing on a dedicated Database Server, such as Microsoft SQL Server. The Repository, Domain, and Service Layers are encapsulated within class libraries compiled into DLLs. These DLLs are referenced by both the Web API and the ML Execution Engine components. The Web API is deployed as a Web Service Application on a Web Server, like Microsoft Internet Information Services (IIS), enabling communication with the Web Client over HTTP. The ML Execution Engine component is deployed as a Windows Service, allowing it to run as a background process on the operating system, performing machine learning tasks independently of user interactions. The Web Client, a web application, is also deployed on IIS, providing users with an interface to interact with the system via a web browser. This deployment architecture ensures scalability, separation of concerns, and efficient communication between the various components of the disclosed system.
[0073] The described architecture emphasizes interoperability among diverse DCI (Data Collection Infrastructure) providers. The flexibility lies in the Repository Layer, which can be replaced with alternative implementations without necessitating significant modifications to the other layers. This modularity allows the system to adapt to different underlying technologies or integration requirements, enhancing its versatility and facilitating future enhancements or migrations. By decoupling the Repository Layer from the rest of the architecture, the system becomes moreresilient to changes, ensuring that it can evolve over time to meet evolving business needs or technological advancements. This design approach fosters adaptability and maintainability, key characteristics of a robust and future-proof software system.
[0074] FIG. 1 illustrates a schematic illustrating the logical sequence of steps in utilizing the present disclosure, alongside the bounded contexts linked with each step. Bounded Contexts serve as distinct conceptual boundaries within the application scope, facilitating clear segmentation and delineation of responsibilities. Furthermore, FIG. 1 highlights the seamless integration of bounded contexts, delineating distinct functional areas within the application scope to enhance clarity and efficiency.
[0075] To get started, users configure templates for Assets, Asset Components, and Predictive Modules, encompassing Predictors and Targets for predictive analytics and continuous monitoring (SI) through functionality embedded within the Digital Templates Bounded Context. Subsequently, users proceed to configure Asset Shapes (S2) using the previously defined asset component and predictive module templates. Upon configuring the Asset Shape, users then utilize functionality within the Asset Information Management Bounded Context to instantiate assets based on the previously defined Asset Shape and map relevant existing tags from the Timeseries Data Historian (TDH) (S3). In S4, users leverage the monitoring strategy auto-discovery features provided within the Machine Learning & Model Management Bounded Context to get automated, data-driven recommendations regarding the optimal monitoring strategy for the predefined Predictive Module Targets. These recommendations encompass suggestions for selecting a combination of SQC and Regression monitoring strategies, alongside the best regression predictors, where applicable, for the defined Targets.
[0076] Upon receiving these recommendations and leveraging domain expertise, users proceed to finalize the Predictive Module Template configuration by implementing the monitoring strategies, utilizing functionalities provided within the Digital Templates Bounded Context (S5). Subsequently, in S6, users utilize functionalities within the Machine Learning & Model Management Bounded Context to batch auto-train and deploy the configured statistical and machine learning models for the instantiated assets.
[0077] As the models are deployed, they are automatically picked up and scored by the ML Execution Engine (S7), operating in the background on a scheduled basis (e.g., every 5 minutes), and the predicted outputs are written to designated tags in the Timeseries Data Historian (TDH), which are already mapped to established attributes in the Contextualization Framework Technology (CFT) server. Importantly, this step occurs independently of user interaction following model deployment. It is noteworthy that statistical models (SQC) and machine learning (ML) model predictions are monitored for abnormal behavior (e.g., deviation from predicted thresholds) within the CFT through its native capabilities, such as rule-based triggered events.
[0078] When abnormal behavior is flagged through continuous monitoring, users in S8 utilize functionalities within the Advanced Data Exploration Bounded Context to analyze the underlying data pertinent to identified anomalous outcomes. In cases of false-positive instances necessitating model retraining, the process loops back to S8, enabling users to manually train the model as needed. If a true positive is identified during data exploration, users initiate the corrective action process (S9). It is important to note that the present disclosure does not incorporate automated physical actionsor any automated equipment control capabilities, but rather supports a human-in-the-loop approach to corrective action processes.
[0079] To ensure clarity in describing the technology layers, methods, and systems in the present disclosure, a realistic use case involving the configuration of digital twins, operationalization, and management of predictive analytics for a Wind Turbine fleet is illustrated throughout the provided diagrams and schematics. This use case serves as a tangible example to demonstrate how the described technologies and processes can be applied in a real-world scenario, offering insight into the practical implementation and benefits of the disclosed systems and methodologies.
[0080] In FIG. 2, a schematic delineates the technology layers and high-level data flow within the Digital Templates Bounded Context. The Template and Asset Shape Repository facilitates CRUD operations on objects intrinsic to the Contextualization Framework Technology (CFT) through a third-party API, transforming them into corresponding Digital Templates Domain Entities showcased in FIG. 3. Furthermore, the Template Service assumes responsibility for generating the suitable type of object and invoking the relevant repository method, in response to requests from the Web Client (Asset Template, Asset Component Template, or Predictive Module Template).
[0081] To illustrate, let's consider a scenario where a user creates an Asset Template for a Wind Turbine, as described in FIG. 4. The exemplary user interface in the Web Client collects user input, constructs a JSON object, and initiates a POST AJAX call to the Template Controller. The Template Controller then binds the JSON object to a Template domain entity and triggers the Create Asset Template function in the Template Service. Notably, in this scenario of creating a new Asset Template, the ID attribute is set to NULL. Furthermore, new Asset Templates utilize the ASSET base template, which has been pre-seeded within the Contextualization Framework Technology (CFT). The ASSET template serves as a foundational template within the CFT, as an empty template designed to designate the Asset concept within the present disclosure. The service proceeds to validate the data integrity of the entity and business rules, returning an error message if validation fails. Upon successful validation, the service invokes the Insert Template function in the Template Repository, which in turn creates the necessary constructs in the CFT, based on the domain entity received from the Template Service. The responsibility of assigning a unique ID (such as a GUID) to the newly created template lies with the CFT. Like the creation of an Asset Template, the process of creating an Asset Component Template, as outlined in FIG. 5, follows a similar pathway. Asset Component Templates utilize the pre-seeded ASSET_COMPONENT template within the CFT. Additionally, the Template Service invokes the same Insert Template function in the Template Repository. This uniform approach is maintained because both Asset Templates and Asset Component Templates share the same underlying native structure, specifically the Element Template, within the CFT.
[0082] In FIG. 6, the exemplary user interface is presented alongside the relevant technology layers involved in the process of Creating or Editing a Predictive Module Template. Like other templates, Predictive Module Templates utilize the pre-seeded PREDICTIVE_MODULE template within the CFT. It's important to note that Predictors and Targets each extend the DataField class depicted in FIG. 3 but possess additional properties. For Predictors, these properties allow configuration as explicit tags, or referencing data from existing attributes within the module's hierarchy (such as from parent components or asset root). Targets require specific configuration regarding the desired monitoring strategy. By default, each target is enabled with the SQC monitoring strategy, but users can alsoenable the Regression monitoring strategy, in which case a list of predictors needs to be specified. Additionally, users can enable Pattern Recognition for the Predictive Module Template, where a list of features (selection from the Predictors and Targets) can be configured. In step SI, users are only required to define the basic information of a Predictive Module Template, along with the list of Predictors and Targets. Subsequently, in step S4, users can leverage monitoring strategy autodiscovery to receive data-driven recommendations on the optimal monitoring strategy for the Targets, and then proceed to step S5 to finalize the Predictive Module Template configuration.
[0083] In FIG. 7, the exemplary user interface is presented alongside the relevant technology layers involved in the process of Creating or Editing an Asset Shape. The Asset Shape functionality enables users to define intricate digital structures for specific asset types by selecting relevant Asset Components and Predictive Module Templates previously configured in the system. For instance, the example illustrates a "Wind Turbine Type A" asset shape tailored for the "Wind Turbine" asset template, configured with a "Rotor" asset component template. Within the Rotor component, a "Blades" module is included, which is based on the "Wind Turbine Rotor Blades Bi-Blade" predictive module template. This depiction showcases the flexibility and customization capabilities afforded by the Asset Shape feature, allowing users to construct complex digital representations of assets to suit their specific requirements and configurations.
[0084] The Asset Shape feature also facilitates the import of 3D scans, e.g., in .gib file format, enabling users to overlay data fields, module predictors, targets, and pattern recognition through drag-and-drop functionality. These 3D scans, e.g., .gib files, are rendered in the Web Client using open-source JavaScript libraries, and each overlapped data item is indexed and tracked by its assigned unique ID. When the user triggers the Save function in the Web Client, the 3D scene is compressed into a base64 string, and a JSON object is constructed, encapsulating the structure of the Asset Shape along with pertinent information. Subsequently, the Web Client initiates a POST AJAX call to the Asset Shape Controller within the Web API. The Asset Shape Controller binds the JSON object to an Assetshape domain entity and triggers the Create Asset Template function in the Asset Shape Service. Notably, in this scenario of creating a new Asset Shape, the ID attribute is set to NULL. The service proceeds to validate the data integrity of the entity and business rules, returning an error message if validation fails. Upon successful validation, the service invokes the Insert Asset Shape function in the Asset Shape Repository, which then creates the necessary constructs in the Contextualization Framework Technology (CFT) leveraging native attribute structures. It also stores the hierarchical relationship of the Asset Shape's template IDs groupings, along with the base64 compression of the 3D scene. It's noteworthy that the responsibility of assigning a unique ID (such as a GUID) to the newly created asset shape also lies with the CFT.
[0085] The present disclosure automates the creation, deployment, and maintenance of Level 3 Digital Twins based on Asset Shapes defined in the Digital Templates, streamlining digitization procedures across various industries and industrial equipment with a practical and efficient solution. The Asset Information Management module enables the collection and management of asset data fields while adjusting the 3D Layout definition defaulted from the Asset Shape.
[0086] Importantly, any changes made to a given Asset Shape are automatically deployed across all instantiated assets in the established DCI, supporting a scalable, iterative process for adding asset components and predictive modules. These powerful concepts are applicable to digitizingmechanical equipment and processes in various industries, but not limited to Oil & Gas, Power Generation, and Manufacturing, which rely on expensive machinery such as, but not limited to compressors, pumps, gas turbines, electric motors, steam turbines, wind turbines, or high voltage generators to ensure optimal performance and production.
[0087] In FIG. 8, a schematic illustrates the technology layers and data flow specifically within the Asset Information Repository Bounded Context. The Asset Repository facilitates CRUD operations on objects intrinsic to the CFT through a third-party API, converting them into corresponding Asset Information Repository Domain Entities, as showcased in FIG. 9. The Asset Service is responsible for validating the data integrity of domain entities and enforcing relevant business rules.
[0088] In FIG. 10A and 10B, the exemplary user interface is presented alongside relevant technology layers involved in the process of Creating or Editing assets. In FIG. 10A, when users select an Asset Shape from the dropdown, a changed event triggers, invoking the Get Blank Asset method in the Asset Controller within the Web API. This request is processed through the Asset Service, which in turn invokes the Assemble Blank Asset method in the Asset Repository. The repository reads the Asset Shape configuration and retrieves the Asset, Asset Component, and Predictive Module Templates as native objects from the Contextualization Framework Technology (CFT).These objects are then translated and data-bound to Asset Information Repository Domain Entities. The resulting blank asset, representing the hierarchical structure of Asset Components, Predictive Modules, and related data fields, including the base64 compression of the 3D scene layout saved in the Asset Shape, without populated data is returned to the Web Client. The Web Client decompresses the 3D scene layout and renders it in the browser. Users are provided with the option to import a more specific 3D scan, e.g., as a .gib file, for the given asset. In such cases, the existing 3D object representing the asset will be replaced while maintaining the relative positions of overlapped data fields.
[0089] In FIG. 10B, users populate the relevant data fields specific to the physical asset being instantiated and save the information by invoking the Save Asset button. A JSON object is constructed in the Web Client, triggering the Save Asset method in the Asset Controller within the Web API through a POST AJAX call. Subsequently, the Save Asset method in the Asset Service validates the request for data integrity and relevant business rules. Upon successful validation, the Insert Asset method is called in the Asset Repository, which creates the necessary constructs in the CFT based on the domain entity received from the Asset Service. It is noteworthy that the responsibility of assigning a unique ID (such as a GUID) to the newly created asset lies with the CFT.
[0090] In FIG. 11A, 11B, and 11C, the exemplary user interface is presented alongside the relevant technology layers involved in the process of Mapping Tags for an Asset. All tag-configured data fields, including Predictors and Targets, for a given instantiated asset are displayed to the users. Users initiate the Smart Tag Map function in the Web Client, which invokes the Smart Tag Map method in the Asset Controller within the Web API. This request is then processed through the Asset Service by the Smart Tag Map method, which extracts pertinent information such as the asset label and the names of the tag-configured attributes defined for the given asset.
[0091] Subsequently, this information is transmitted to the Asset Repository by invoking the Tag Keyword Match method. Here, a wildcard tag search is conducted in the established Timeseries Data Historian (TDH) based on the asset label (e.g., searching for tags containing 'WT 001' or variationssuch as 'WT_001' or 'WT.001', which are common naming conventions in industrial applications). Fuzzy logic is applied to the resulting set of tags for each attribute name. Some popular fuzzy logic search algorithms that may be employed include Levenshtein distance or Jaro-Winkler distance. For instance, the attribute named 'Acceleration' may be cross-referenced to an instrument tag named 'wt.OOl.accel' with a confidence level of 57% using such fuzzy logic search algorithms. These results are then returned to the Web Client, as depicted in FIG. 11A. If the auto-tag mapping is unsuccessful or provides low confidence, users can opt to perform a manual tag search, as depicted in FIG. 11B. Once users have confirmed the tag mapping, they can apply it to the established CFT, as shown in FIG. 11C.
[0092] In FIG. 12, the exemplary user interface represents a fully implemented Asset upon completion of Step S3, alongside the relevant technology layers responsible for accomplishing this task. The Read Asset method in the Asset Repository is invoked, extracting snapshots for all relevant attributes configured for the given asset. These snapshots are then data-bound to the values of the respective attributes, and the domain entity is returned to the Web Client. Subsequently, the 3D scene is re-rendered to update the values of overlapped data fields, providing users with an updated and comprehensive view of the asset's data status.
[0093] Upon completion of Step S3, the digital structure of the asset, along with relevant modules and targeted data for continuous monitoring through predictive analytics, is defined in the Asset Shape. Additionally, assets are instantiated based on the Asset Shape, with relevant data fields populated and instrument tags mapped. Moving forward to Step S4, monitoring strategy autodiscovery features are utilized to obtain data-driven recommendations regarding the optimal machine learning models and predictors (where applicable) for the defined targets within any predictive module. These features are provided within the Machine Learning & Model Management Bounded Context, which also offers comprehensive oversight of SQC and Machine Learning models coverage, accuracy, batch training, and deployment features that support hyper scalability of the system.
[0094] FIG. 13 presents a schematic illustrating the technology layers and data flow specifically within the Machine Learning & Model Management Bounded Context. Within this context, domain entities such as ModelBase, ModeISQC, ModelRegression, Modelclustering, and ModelAssignment as depicted in FIG. 14A, contain classes for managing statistical and machine learning models along with their assignments to assets. CRUD operations on these entities are performed by the ML Model Persistent Repository, with the persistence system being a dedicated SQL Database.
[0095] FIG. 15 showcases the relational database diagram featuring pertinent tables for machine learning models and their assignments to assets. The communication between the ML Model Persistent Repository and the SQL Database is facilitated through an Object Relationship Model (ORM) library such as Entity Framework. Additionally, the CFT seamlessly links to the dedicated SQL Database via an OLEDB driver, enabling efficient table lookup of machine learning models and relevant metadata into CFT attributes.
[0096] Furthermore, the SQC, Regression, Regressionstrategy, Clustering, and Cluster domain entities, as depicted in FIG. 14A, consist of classes with attributes and logic designed for executing the training and prediction of machine learning models. These entities encapsulate open source machine learning libraries, exposing methods for fitting, predicting, batch predicting, and serializingmachine learning objects. The ML Model Operations Service utilizes these entities to facilitate the actual execution of machine learning operations, providing a cohesive interface for training and predicting models within the system.
[0097] The ModelManagementSummary, ModelManagementAssetSummary, Model ManagementSummaryOverallCoverage, and ModelManagementSummaryComplexityDistribution domain entities outlined in FIG. 14B are leveraged by the ML Model Management Repository to generate summaries pertaining to the coverage and accuracy of statistical and machine learning models. These entities enable the repository to aggregate and analyze data related to model performance, providing insights into the overall effectiveness and complexity distribution of the models deployed within the system. The Data ETL (Extraction, Transformation, and Loading) Domain Entities presented in FIG. 14B encompass classes responsible for storing both historical and snapshot timeseries data. These entities are utilized by the Data ETL Repository to extract data from the Contextualization Framework Technology (CFT), subsequently performing transformation and loading operations. These operations are crucial for constructing and data-binding the domain entities, enabling efficient handling and utilization of timeseries data within the system.
[0098] In FIG. 16, the exemplary user interface is presented alongside the relevant technology layers involved in the process of Monitoring Strategy Auto-Discovery. Users begin by selecting an asset shape and a predictive module template for analysis. A list of instantiated assets based on the selected asset shape is then populated in the user interface. Upon selecting the desired instantiated assets to include in the auto-discovery analysis and desired timeframe, users invoke the Analyze button. Subsequently, the Web Client triggers the Strategy Auto-Discovery method in the ML Model Operations Controller within the Web API via an AJAX POST request. The Data ETL Repository extracts historical timeseries data for all predictors and targets configured in the given predictive module template for all selected instantiated assets, while they have been operating at steady state. This dataset is then provided to the ML Model Operations Service via the Data ETL Service. The Discover Monitoring Strategy method within the ML Model Operations Service processes the dataset according to the flow chart depicted in FIG. 17.
[0099] First, all asset's data is combined and cleansed by removing outliers using quartile methodology. For each target within a predictive module, the coefficient of variance (CoV) for the combined dataset is calculated using the following formula:Dataset Standard Deviatian100Dataset' Mean
[0100] The coefficient of variation (CoV) is a statistical measure that expresses the relative variability of a set of data points in comparison to the mean (average) of the data set. It is often used to assess the degree of dispersion associated with a particular dataset.
[0101] Interpretation of the CoV: o A low coefficient of variation suggests that the data points are closely clustered around the mean, indicating low relative variability.o A high coefficient of variation suggests that the data points are more spread out from the mean, indicating higher relative variability.
[0102] If the coefficient of variation (CoV) for the specified target exceeds 50%, it is advisable to implement a Regression monitoring strategy alongside the Statistical Quality Control (SQC) methodology. This recommendation is made to address the heightened variability in the target, which necessitates the exploration of additional predictors.
[0103] To identify the optimal predictors for the given target, the analysis employs the Ordinary Least Squares (OLS) regression algorithm. This algorithm efficiently identifies linear relationships between variables in the dataset. The Root Mean Squared Error (RMSE) accuracy metric is then calculated for the OLS model. Subsequently, a Gradient Boosting Machine (GBM) regression algorithm is applied to the same dataset. GBM, being a higher complexity algorithm, can capture non-linear relationships between variables. The RMSE accuracy metric is computed for the GBM model. If the GBM algorithm exhibits an accuracy 1% or more superior to that of OLS, the GBM fitted model is considered for identifying the optimal predictors for the given target.
[0104] To facilitate this process, the analysis employs the Permutation Feature Importance (PFI) methodology. PFI assesses the impact of each predictor on the model's accuracy by measuring the percentage change in RMSE when a specific predictor is removed from the dataset. Predictors with a positive impact are deemed important for prediction.
[0105] The output of the Discover Monitoring Strategy method, containing the monitoring strategy recommendations, is sent as a response to the Web Client and rendered in the user interface, as depicted in FIG. 16.
[0106] Although experts may possess an intuitive understanding of suitable Regression monitoring strategies and adequate predictors, the Auto-Discovery feature serves as an additional layer of support. This feature not only validates domain experts' intuitions but also provides valuable insights by uncovering nuanced correlations that may remain unnoticed, even with years of domain expertise. By leveraging data-driven recommendations, the system contributes to a more comprehensive and accurate configuration of monitoring strategies. Furthermore, the Monitoring Strategy AutoDiscovery feature establishes a sandboxing environment. In this controlled space, domain experts can conduct research and experimentation to test new strategies for continuous monitoring and improvement. This functionality empowers experts to explore innovative approaches, fostering a dynamic and responsive system that adapts to evolving industrial processes.
[0107] Based on data-driven recommendations and domain knowledge, users proceed to finalize the predictive module template configuration (Step S5) by implementing monitoring strategies for all defined targets. FIG. 18 illustrates the exemplary user interface alongside relevant technology layers involved in configuring Regression and Pattern Recognition within a Predictive Module Template. For multi-variate regression analytics, users configure the list of predictor variables (predictors) that explain the normal variation of the given target, along with the evaluation rate and time outside the prediction threshold. Predictor variables are selected from the list of previously defined Predictors and / or other Targets. Users can also enable Pattern Recognition for the given Predictive Module by selecting relevant features based on domain knowledge. Features are chosen from the list of predictors and targets configured in the predictive module template. Additionally,users need to set the evaluation rate and a duration threshold for pattern recognition to ensure accurate analysis and prediction. Finally, users save the final configuration for the predictive module template by invoking the Save button. A JSON object is constructed in the Web Client, triggering the Update Predictive Module Template method in the Template Controller within the Web API through a PUT AJAX call. Subsequently, the Update Predictive Module Template method in the Template Service validates the request for data integrity and relevant business rules. Upon successful validation, the Update Template method is called in the Template Repository, which creates the necessary constructs in the CFT based on the domain entity received from the Template Service.
[0108] Scaling up the operationalization of machine learning models presents a substantial challenge due to the diverse landscape of asset makes, models, and vintages across various industries. Additionally, differing equipment maintenance practices employed by organizations compound this complexity, requiring predictive models tailored to each asset or industrial process to accurately capture normal behavior. Furthermore, the challenge intensifies when equipment components undergo repair or replacement, as retraining models with limited historical data becomes an additional obstacle, impacting model accuracy and effectiveness. Existing digital twin solutions often lack practicality and scalability, failing to strike a suitable balance of nimbleness, complexity, and automation in operationalizing and maintaining predictive analytics models.
[0109] To address these challenges, the present disclosure introduces methods enabling autonomous model training in batch and efficient model management. Users can batch train SQC and Regression models for an entire asset, offering scalability when digitizing new assets. FIG. 20 showcases the exemplary user interface alongside relevant technology layers involved in the batch auto-training process of statistical and machine learning models configured for a given instantiated asset. Upon selecting a desired training timeframe for the asset, users trigger the Auto-Train button in the Web Client, invoking the Auto-Train method in the ML Model Management Controller within the Web API. The Data ETL Repository's Read Asset Historical Data method extracts timeseries historical data for all predictors and targets configured across predictive modules for the selected asset during the specified timeframe. This historical training dataset is then passed to the ML Model Operations Service via the Data ETL Service. Subsequently, the Auto-Train Asset method in the ML Model Management Service, in collaboration with the Batch Train Predictive Module method in the ML Model Operations Service, processes the historical dataset according to the process flow diagram depicted in FIG 21A, 21B, 21C, and 21D. This streamlined process enables efficient batch auto- training of models, facilitating scalability and enhancing the effectiveness of predictive analytics within the system.
[0110] For all Predictive Module Targets, Statistical Quality Control (SQC) Models are trained as the base layer of monitoring complexity. SQC continuous monitoring strategy is scalable because it uses computationally efficient statistical methods and can be easily automated to monitor small, medium, and large datasets, reducing the need for manual intervention. This method is used as the primary line of defense when limited historical data is available for a given asset or asset component (i.e., newly installed asset, or part replacement). Upper and lower control limits are calculated based on the Mean, Standard Deviation, and a default Outlier Threshold Factor (OTF), which can be manually adjusted by users after the SQC model is trained.
[0111] For Targets that have Regression configured in the Predictive Module Definition a regression algorithm is trained. Regression is a machine learning method used to predict the value of a dependent variable (predicted variable) based on the values of one or more independent variables (predictors). The present disclosure automatically extracts historical data for the given Target (predicted variable) and all predictor variables when the asset or industrial process was operating at steady state, removes outliers across the entire dataset using quartile logic, and splits the data set into training (80%) and testing (20%) sets. The split is performed randomly. Note that if any predictor variables are categorical, stratification is applied to the split process to preserve categorical states original distribution.
[0112] Next, various types of regression algorithms are fitted through the training set, starting with Ordinary Least Squares (OLS) and then moving to more complex algorithms such as Prediction Tree, Random Forest, or Gradient Boosting Machine (GBM). Each algorithm is automatically trained on the training set and tested on the testing set, respectively, and trained on the overall historical data set. Through this process, three different root mean squared errors (RMSE) are computed for each algorithm type: Training RMSE, Testing RMSE, and Overall RMSE, along with complexity measures such as the Model Size (bytes) and Training Speed (seconds). To perform autonomous algorithm selection, a Balance Metric is computed through a weighted additive equation:
[0113] The default weights are set as follows (but can be manually adjusted by users):
[0114] Over-fit validation is also performed for each algorithm type by analyzing the ratio between the Training RMSE and Testing RMSE. A Predictability Factor is computed as follows:Pretiict &Oifv actor
[0115] A predictability factor near the value 1 indicates good predictability, while a factor greater than 2 suggests the given algorithm may be overfitting the data. The optimal regression algorithm is selected as the model with highest balance score that does not present an overfit concern.
[0116] Prediction Intervals are computed based on the Overall RMSE and a default Outlier Threshold Factor (OTF), which can be manually adjusted by users after the model is trained.
[0117] The Normal Variation Range (NVR) is calculated as an accuracy measure for SQC and Regression models.
[0118] For SQC models, NVR is computed as follows:NVR = 2 * OTF * Standard Deviation
[0119] For Regression models, NVR is computed as follows:NVR = 2 * OTF * Overall RMSE
[0120] Furthermore, the Fleet Average Normal Variation Range (FANVR) for SQC and Regression models is computed. This is the average across all existing models trained for Assets implementing a given Asset Shape and a given Predictive Module Target.
[0121] If pattern recognition is enabled in the Predictive Module Template, a HDBSCAN clustering algorithm is trained. The HDBSCAN is a machine learning algorithm used to partition a multi-variate dataset into clusters based on the similarity of the data points. Note that with this method does not attempt to predict any given variable, but rather discover patterns in a multivariate data space.
[0122] The present disclosure automatically extracts historical data for all the feature variables (Predictors and Targets) configured in the Predictive Module Template when the asset was operating at steady state and removes outliers across the entire dataset using quartile logic. To auto-discover the optimal number of clusters (patterns) in the training dataset, a list of minimum cluster sizes (minimum amount of data points in a cluster as percent of the overall training data set) and a list of distance functions (mathematical functions to measure the distance between two data points; examples include Euclidian distance, cosine distance, etc.) are provided. For each combination of minimum cluster size and distance function, separate HDBSCAN cluster fits are computed. The optimal fit is selected by using the silhouette method, picking the fit with the highest silhouette score.
[0123] Given that the HDBSCAN is accepted in the literature as a state-of-the-art algorithm for unsupervised cluster discovery, scoring new data points (observations) with a pre-fitted HDBSCAN is challenging mathematically (given a new observation, predict which cluster it most likely belongs to). Although HDBSCAN provides this capability through a native implementation, based on our testing, the native scoring algorithm leaves room for improvement in accuracy. To surpass this challenge, the present disclosure fits a Multivariate Empirical Distribution through each cluster that was identified through the HDBSCAN fitting process. The distributions are saved along with each cluster, providing a more accurate mechanism to score new observations. To score a new observation, the likelihood for each cluster is computed using the respective cluster empirical distribution Probability Density Function (PDF). The most likely cluster is selected based on the highest PDF. The highest PDF is also returned as the Overall Score of the new observation. The anomaly detection mechanism works by computing the Overall Score of the HDBSCAN over the training data set and setting up lower and upper control limits based on the Mean, Standard Deviation of the overall score, and a default Outlier Threshold Factor (OTF), which can also be manually adjusted by users. As new observations (real-time readings of all configured feature variables) are evaluated through the model, the Overall Score of the new observation is computed, along with identifying the most likely cluster it belongs to. An anomaly is flagged if the Overall Score of the new observation is outside the control limits (unclassified), or if the most likely cluster has been labeled as abnormal by the user in the training phase. For an unclassified observation, the effect of each feature is computed by replacing the reading of a given feature iteratively with each cluster's Mean for that feature and recalculating the Overall Score, measuring the percent change in the Overall Score.
[0124] The auto-training results are transmitted to the Web Client and displayed in the user interface, as illustrated in FIG. 20. However, users retain the final decision-making authority to deploy an auto-trained model. FIG. 22 showcases the exemplary user interface alongside the relevant technology layers involved in the process of deploying auto-trained predictive models for a given asset.
[0125] If a statistical or regression model meets accuracy expectations, users have the option to deploy it by invoking the Deploy button for the respective model. This action triggers the Save Model method in the ML Model Persistence Controller within the Web API through an AJAX POST call. Subsequently, the Create Model method in the ML Model Persistence Service is invoked, which is accomplished by subsequent calls to the Insert Model and Insert Asset Model Assignment methods in the ML Model Persistence Repository.
[0126] If a statistical or regression model fails to meet accuracy expectations, users have the option to inspect the model and perform manual training. FIG. 23 illustrates the exemplary user interface alongside the relevant technology layers involved in the process of manually training an SQC model for a given target. Similarly, FIG. 24A and 24B depict the exemplary user interface alongside the pertinent technology layers involved in the process of manually training a Regression model for a given target. These interfaces provide users with the necessary tools and functionalities to adjust and refine the models manually, ensuring optimal performance and accuracy according to specific requirements and domain expertise.
[0127] It's important to note that the present disclosure necessitates user inspection of Pattern Recognition models before deployment, given the requirement for human intervention in this scenario. FIG. 25A and 25B provide a visual representation of the exemplary user interface alongside the relevant technology layers involved in the process of manually training a Clustering model for a given Predictive Module. These interfaces offer users the capability to review and fine-tune Clustering models manually along with labeling discovered clusters using domain knowledge, ensuring alignment with specific requirements and domain expertise before deployment.
[0128] As an increasing number of models are deployed, it becomes essential to have visibility into their coverage and accuracy across the entire asset fleet. The present disclosure addresses this need with the Model Coverage dashboard, which provides high-level metrics on model coverage, gaps, and complexity distribution. Additionally, the present disclosure automatically identifies outliers within the fleet for both SQC and Regression models, ensuring accurate monitoring of all assets. To accomplish this, all existing models are grouped by asset shape and the group mean and standard deviation of the Normal Variation Ranges are computed. Models that exceed the group mean plus three times the standard deviation are flagged as fleet outliers. Furthermore, the system automatically detects Targets and Pattern Recognition that lack trained models. FIG. 19 illustrates the exemplary user interface along with the relevant technology layers utilized for generating the Model Management Summary Dashboard.
[0129] Upon completing Step S6, where predictive models are trained and deployed for instantiated assets, the present disclosure provides a mechanism for continuous scoring and monitoring of these deployed models in near real-time (e.g., every 5 minutes) in Step S7. This is facilitated through the ML Execution Engine, as illustrated in FIG. 26, which operates as a scheduledtask (Windows Service) and utilizes various services and repositories including ML Model Management, Data ETL, and ML Model Operation Services.
[0130] The ML Model Management Repository extracts configuration details of all predictive modules for instantiated assets from the Contextualization Framework Technology (CFT), accessing predictive model information via a native OLEDB connection to the dedicated SQL Database. Concurrently, the Data ETL Repository retrieves the latest sensor readings (snapshots) for targets and predictors based on the extracted predictive module configurations.
[0131] The ML Execution Engine processes this information by splitting predictive modules into groups and initiating background threads for each group, leveraging deserialized predictive models to perform predictions for regression and clustering models. These predictions are then written back to designated tags in the Timeseries Data Historian (TDH) through the Data ETL Service and Repository, mapped to attributes within the CFT.
[0132] Continuous monitoring within the CFT involves triggering events based on predefined rules. Note that this is native functionality within the CFT. For SQC models, events are generated if the target snapshot value falls outside control limits calculated using historical mean, standard deviation, and user-defined outlier threshold factor (OTF).SQC Upper Control Limit = Historical Mean + OTF * Historical Standard DeviationSQC Lower Control Limit = Historical Mean - OTF * Historical Standard Deviation
[0133] Similarly, for Regression models, events are triggered if the target snapshot falls outside the prediction interval based on model RMSE and OTF.RGR Upper Prediction = Predicted Snapshot + OTF * Model RMSERGR Lower Prediction = Predicted Snapshot - OTF * Model RMSE
[0134] For Pattern Recognition Clustering models, events are triggered if the predicted overall score snapshot is outside historically trained score control limits or if the predicted cluster is labeled as abnormal (cluster labeling was done by the user in the training phase).
[0135] Efficiently managing anomalous outcomes is crucial, especially with many deployed predictive models and real-time continuous monitoring. The Active Anomaly dashboard provides a near real-time display of active anomalous outcomes. The Rate of Change (RoC) is calculated by fitting a straight line to the last 12 hours of sampled data for the Target and computing the slope, indicating the direction and steepness of the abnormal trend. It's important to note that RoC doesn't apply to Pattern Recognition. Anomalous targets are grouped by the Asset.
[0136] FIG. 29 illustrates the exemplary user interface alongside the relevant technology layers utilized for generating the Active Anomalies Dashboard. The exemplary user interface triggers the Get Active Anomalies method in the Data Exploration Controller via a Javascript timer through a GET AJAX call. The Read Active Anomalies Method in the Data Exploration Repository extracts all open events along with historical data in the past 12 hours from the Contextualization Framework Technology (CFT) and transmits the information to the Data Exploration Controller via the Data Exploration Service. The Service computes RoC and instantiates the Data Exploration Domain Entitiesdepicted in FIG. 27 and FIG. 28. Finally, the information containing the relevant domain entities is transmitted to the Web Client and rendered by the user interface.
[0137] Users can drill down into any asset summary to access near real-time information about all configured predictive module anomaly states, including the asset 3D model overlaid with sensor readings. This dynamic Asset Anomaly Summary screen is generated based on the Asset Shape definition and data configured in the Asset Information Management. The 3D model overlay functionality offers a powerful tool that provides users with a comprehensive view of equipment's real-time status, facilitating quick identification and resolution of any issues.
[0138] FIG. 30 illustrates the exemplary user interface alongside the relevant technology layers utilized for generating the Asset Anomaly Summary Dashboard. The exemplary user interface initiates the Get Asset Summary method in the Data Exploration Controller through a Javascript timer using a GET AJAX call. The Read Asset Summary Method in the Data Exploration Repository retrieves snapshots of predictive module anomalous states, along with the serialized 3D scene and referenced overlapped data field snapshots. This information is then processed by the Data Exploration Service, where Data Exploration Domain Entities depicted in FIG. 28 are instantiated. Finally, the relevant domain entities containing the information are transmitted to the Web Client and rendered by the user interface.
[0139] To provide a comprehensive analysis of any active anomaly, the disclosed Advanced Data Exploration tool offers users access to detailed insights. Users can access this tool by browsing the predictive module list or by clicking on any overlaid data on the 3D model. The Advanced Data Exploration tool streamlines the anomaly investigation process by providing a comprehensive set of data exploration tools. With this tool, users can quickly delve into the underlying data causing any anomaly, ensuring transparency and clarity in understanding the root cause of an issue.
[0140] FIG. 31 illustrates the exemplary user interface alongside the relevant technology layers involved in the process of advanced data exploration (Step S8) for a given Predictive Module. This exploration is facilitated through the Predictive Models dashboard and the Temporal Data Analysis functionality.
[0141] The user initiates the exploration by triggering the Refresh Predictive Models Data and Load Temporal Data buttons in the Web Client. These actions invoke respective methods in the Data Exploration Controller within the Web API through a GET AJAX call. The Read Predictive Models Data method in the Data Exploration Repository extracts snapshot data and model metadata from the CFT. Simultaneously, the Read Predictive Module Historical Data method in the Data ETL Repository extracts historical timeseries data for all predictors and targets configured in the given predictive module for the user-defined timeframe.
[0142] This information is transmitted to the Data Exploration Controller through the Data Exploration and Data ETL Service, where Data Exploration Domain Entities depicted in FIG. 27 and FIG. 28 are instantiated. Finally, the relevant domain entities containing the information are transmitted to the Web Client and rendered by the user interface, enabling users to conduct advanced data exploration and gain insights into anomaly root causes.
[0143] The predictors and targets Timeseries Trends depicted in FIG. 31 are meticulously designed to reveal the long-term, continuous behavior of all relevant data within the module. To achieve this, the series are grouped according to their order of magnitude, and each group is assigned an appropriate Y Axis Scale. The Y Axis Scale is determined by computing the average of each series and expressing it in power of 10. Series that share the same power of 10 are assigned the same Y Axis Scale, with the Min and Max values automatically computed based on the absolute minimum and maximum of all series in the group. A 5% scale factor is then applied to the Scale Min and Max to ensure appropriate default series overlapping and zooming on both trends. Users do have the option to customize the Y Scale series assignments and Min / Max values using the Trend Setting feature. Moreover, users can also create and chain multiple conditional filters to further filter the data on the trends in case a specific operational profile is required for the analysis. To delve deeper into the anomaly, the present disclosure provides the Bivariate Analysis feature, helping users to gain insight into the time progression of the relationship between any two variables in the predictive module. By selecting a combination of targets and / or predictors, the exemplary user interface generates a scatter plot and colors the data points based on various time intervals (weekly, monthly, or yearly). This feature can help understand subtle outliers that may not be apparent in the timeseries trends.
[0144] Furthermore, leveraging the Asset Shape and modular design concepts, the present disclosure offers Fleet Comparison capabilities, enabling users to easily compare relevant data across all assets that are based on the same Asset Shape (peer equipment analysis) further consolidating their analysis. For Targets that are configured with SQC monitoring strategy, the user interface automatically generates bar charts, providing a comparison of descriptive statistics across peer equipment. Moreover, bivariate peer fleet comparison functionality is disclosed, whereby users can select a combination of targets and / or predictors, along with a specific time frame, to generate a scatter plot that colors the data points based on the asset they belong to. Finally, the Advanced Data Exploration tool offers functionality to retrain statistical and machine learning models on the fly (Loop back to Step S6). This ensures that the models used to analyze the data are always up-to-date and optimized for the specific use case. This not only saves time, but it also improves the accuracy of the continuous monitoring process.
[0145] FIG. 32a and 32b depict the exemplary user interface alongside pertinent technology layers involved in the process of advanced data exploration (Step S8) for a given Predictive Module.
Claims
AMENDED CLAIMS received by the International Bureau on 03 October 2025 (03.10.2025)1. A computer-implemented method for deploying and operating digital twins of aw industrial assets, the method comprising:(a) accessing, via a native application programming interface (API), a pre-existing contextualization framework of a data collection infrastructure (DCI) that maintains historian tags, asset definitions, and event rules for physical industrial assets;(b) defining an Asset Shape as a reusable meta-template that:(i) specifies a hierarchical arrangement of asset components and predictive modules,(ii) pre-associates each predictive module with one or more predictors, targets, and monitoring strategy definitions,(iii) includes tag-mapping logic for historian tags of the DCI, and(iv) optionally binds data fields to positional elements in a three-dimensional layout of the asset;(c) instantiating, from the Asset Shape, a plurality of asset instances by automatically generating native objects within the contextualization framework and applying the tag-mapping logic to associate predictors and targets with corresponding historian tags;(d) performing a monitoring strategy auto-discovery process on historical data retrieved from the DCI to:(i) computing variability measures for each target,(ii) selecting a tiered monitoring strategy including at least one of:(1) a univariate statistical quality control (SQC) model,(2) a multivariate regression model with automatic predictor selection, and(3) a clustering-based pattern recognition model, and(iii) for regression, selecting an algorithm by applying a balance metric combining model accuracy and complexity;(e) for clustering-based pattern recognition, generating clusters using a density-based algorithm and, for each cluster, fitting a multivariate empirical probability distribution, wherein scoring a new observation comprises computing likelihoods under each cluster distribution, selecting a most likely cluster, and detecting an anomaly if the likelihood is outside a statistical control limit or if the cluster is pre-labeled abnormal;(f) deploying the selected models for the plurality of asset instances;(g) on a scheduled basis, executing the deployed models by:28(i) retrieving latest predictor values from the DCI,(ii) generating predictions or anomaly scores, and(iii) writing the predictions or scores back into native historian tags of the DCI; and(h) allowing a user, via an anomaly investigation interface, to view model outputs overlaid on the three-dimensional layout, and to initiate in-place retraining and redeployment of a model without re-authoring the asset definition.
2. The method of claim 1, wherein the variability measures comprise a coefficient of variation, and regression models are selected for targets whose coefficient exceeds a threshold.
3. The method of claim 1, wherein the automatic predictor selection is used for regression using permutation feature importance to identify predictors with positive impact on model accuracy.
4. The method of claim 1, wherein the balance metric comprises a weighted combination of reciprocal overall root-mean-squared error and reciprocal model size, with user- adjustable weights.
5. The method of claim 1, wherein the SQC model computes upper and lower control limits as mean ± k-o, with k being a configurable outlier threshold factor.
6. The method of claim 1, wherein the deployed selected models are grouped by Asset Shape for fleet-level coverage analysis, and outlier models are identified by normal variation range exceeding mean ± m-o across the fleet.
7. The method of claim 1, wherein the three-dimensional layout is stored as a compressed scene file with bindings of data fields to positional identifiers and is deployed in synchrony with the model instantiation.
8. The method of claim 1, wherein the anomaly investigation interface includes bivariate analysis and peer-fleet comparison views for predictor and target data.
9. The method of claim 1, wherein the in-place retraining of the model occurs via the anomaly investigation interface and triggers an update of the corresponding serialized model in internal storage and immediate redeployment to the execution engine.
10. The method of claim 1, wherein the writing the predictions or scores into the native historian tags of the DCI enables the contextualization framework's native event rules to generate alarms without separate alarm infrastructure.
11. A predictive digital twin system for operationalizing statistical and machine learning models within a pre-existing data collection infrastructure (DCI), the system comprising:(a) a repository layer configured to read and write native objects to a contextualization framework of the DCI via a native API, including historian tag mappings, asset definitions, and event rules;(b) an internal storage database storing serialized predictive model objects and asset-to- model assignments;(c) a service layer executing business logic to:(i) create an Asset Shape meta-template,(ii) instantiate assets from the Asset Shape by generating native contextualization framework objects,(iii) perform monitoring strategy auto-discovery,(iv) train models using historical data from the DCI, including empirical distribution scoring for clustering models, and(v) deploy trained models for execution;(d) a machine learning execution engine configured to run on a schedule to:(i) retrieve deployed models from the internal storage,(ii) query latest predictor values from the DCI,(iii) produce predictions or anomaly scores, and(iv) write results into historian tags of the DCI for native event rule processing; and(e) a user interface configured to display three-dimensional asset layouts with overlaid predictor and target data, anomaly states, and to initiate in-place retraining and redeployment of models without modifying the Asset Shape.
12. The system of claim 11, wherein the repository layer is replaceable to support multiple vendor-specific contextualization frameworks without modifying the service layer.
13. The system of claim 11, wherein the machine learning execution engine deserializes models into memory, processes them in parallel threads, and writes results into historian tags via the contextualization framework API.
14. The system of claim 11, wherein the service layer applies a progressive complexity approach in which SQC models are trained for all targets, regression models are added when variability thresholds are exceeded, and pattern recognition is optionally added based on configuration or auto-discovery results.
15. The system of claim 11, wherein the user interface is configured to dynamically render anomaly states on the three-dimensional layout using real-time tag values from the DCI.
16. The system of claim 11, further comprising fleet-level dashboards aggregate model coverage and performance statistics across all assets sharing an Asset Shape, compute average normal variation ranges, and flag outlier models.
17. The system of claim 11, further comprising clustering-based pattern recognition models are generated using a density-based algorithm and scored using multivariate empirical probability distributions.
18. The system of claim 11, wherein the monitoring strategy auto-discovery uses historical data from multiple assets of the same Asset Shape to recommend predictors and algorithms.
19. The system of claim 11, further comprising an object-relational mapping interface via which the internal storage database is accessible to both the service layer and the execution engine.
20. The system of claim 11, wherein the user interface comprises controls to adjust model parameters, including outlier threshold factors, prediction intervals, and variability thresholds, without redeploying the Asset Shape.