Industrial Plant ML Abstraction Layer for Standardized Data Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Connecting machine learning to industrial plant data is challenging due to data distribution across multiple systems, high dimensionality, and the need for extensive configuration and re-engineering, which leads to overfitting and high costs in data labeling and model deployment, especially in process control and automation where each plant has unique automation systems and sensors.
Innovation Solution
An industrial plant machine learning system with an abstraction layer using a machine learning markup language for standardized communication between the machine learning model and the industrial plant, allowing for automatic data extraction and labeling, reducing configuration efforts, and enabling secure, structured access to plant data, thus facilitating model deployment and transfer across plants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine learning models are trained on industrial plant data, then model performance improves, but data distribution across multiple systems increases complexity
Solution Approach 1:
The patent introduces a data lake as an intermediary layer between distributed industrial plant data systems and machine learning models. This data lake consolidates data from multiple sources (DCS, SIS, LIMS, ERP, CMMS) into a unified repository, simplifying access while maintaining the distributed nature of source systems. The intermediary layer handles data integration, transformation, and standardization, allowing ML models to access comprehensive plant data without directly managing the complexity of multiple distributed systems.
2Ease of operation
If extensive configuration is performed to connect ML models to plant data, then data access improves, but re-engineering efforts increase
Solution Approach 1:
The patent implements preliminary action by pre-configuring the data lake with standardized data collection frameworks, integration patterns, and access protocols before ML model deployment. Data pipelines are established in advance from various plant systems to the data lake, with predefined data transformation rules and quality assurance mechanisms. This preliminary setup eliminates the need for extensive re-engineering when deploying new ML models, as the infrastructure is already in place to support diverse data access requirements.
3Measurement precision
If high dimensionality of industrial data is used, then model accuracy improves, but overfitting risk increases
Solution Approach 1:
The patent applies the extraction principle by selectively extracting and filtering relevant features from high-dimensional industrial data before feeding them to ML models. The data lake implements data preprocessing pipelines that identify and remove redundant, noisy, or irrelevant data points while preserving critical predictive features. This extraction process reduces dimensionality while maintaining the essential information needed for accurate predictions, thereby preventing overfitting.
4Measurement precision
If data labeling is performed manually, then model training quality improves, but costs increase
Solution Approach 1:
The patent implements self-service through automated data labeling mechanisms within the data lake framework. The system uses pre-configured labeling rules, automated annotation tools, and intelligent algorithms to label data without requiring extensive manual human intervention. This self-service approach maintains training quality by applying consistent, reproducible labeling standards while dramatically reducing the labor costs associated with manual data annotation.
Data Source
AI summary
An industrial plant machine learning system includes a machine learning model, providing machine learning data, an industrial plant providing plant data and an abstraction layer, connecting the machine learning model and the industrial plant, wherein the abstraction layer is configured to provide standardized communication between the machine learning model and the industrial plant, using a machine learning markup language.
