Digital Twin Database Identifier Alignment via Autoencoder
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Digital twin databases face inconsistencies due to discrepancies in naming standards across heterogeneous data sources, leading to incomplete and inaccurate information, which requires manual correction by domain experts, resulting in high costs and time consumption.
Innovation Solution
A method and system that utilize an encoder to compute latent representations of equipment identifiers from different data sources, compare these representations using a similarity metric, and update the digital twin database by aligning identifiers with a similarity score above a threshold, thereby automating the process of maintaining data consistency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review and correction by domain experts is used to reconcile identifier discrepancies, then data accuracy is improved, but cost and time consumption increase significantly
Solution Approach 1:
The system enables self-service by automatically reconciling identifier discrepancies through AI-based similarity comparison. The encoder computes latent representations of identifiers from different data sources, the similarity metric compares these representations, and the updating component automatically aligns identifiers without requiring manual intervention from domain experts, thus resolving the contradiction between accuracy and time consumption
Solution Approach 2:
The patent replaces the mechanical manual review process with an automated AI-based system. Instead of human experts manually comparing and correcting identifier discrepancies, the system uses encoders, similarity metrics, and automated updating mechanisms to perform the same function, significantly reducing time consumption while maintaining data accuracy
2Reliability
If manual creation of matching identifier tables is used during onboarding, then data consistency is improved, but productivity decreases due to expensive and time-consuming processes
Solution Approach 1:
The system performs preliminary action by pre-training the encoder on diverse identifier formats and naming standards before actual onboarding occurs. This pre-training enables the encoder to immediately compute accurate latent representations and similarity scores when new data sources are onboarded, eliminating the need for manual creation of matching tables during the onboarding process itself
Solution Approach 2:
The automated system enables self-service during onboarding by automatically reconciling identifiers from new data sources. The encoder computes latent representations, the similarity metric compares identifiers, and the updating component automatically aligns them, eliminating the need for manual expert intervention and significantly improving onboarding productivity while maintaining data consistency
3Productivity
If automated AI-based alignment is used to reconcile identifiers, then productivity is improved, but measurement precision may be compromised due to similarity score thresholds
Solution Approach 1:
The system implements feedback mechanisms where the similarity metric continuously compares latent representations and provides feedback on matching confidence. The updating component uses this feedback to adjust alignment decisions, ensuring that only identifiers with sufficient similarity scores are automatically aligned, thus maintaining measurement precision while achieving high productivity
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting similarity score thresholds and encoder parameters based on the specific data sources and contexts. This allows the system to optimize both productivity and measurement precision by adapting the alignment criteria to the particular characteristics of each data source, ensuring accurate matching while maintaining fast processing speeds
Data Source
AI summary
To restore consistency of a digital twin database, identifiers with metadata imported from various data sources are processed by an encoder, which computes latent representations of the identifiers that are compared by an efficient similarity metric. If the respective similarity score exceeds a threshold, a match is detected between the identifiers. In that case, the digital twin database is updated by aligning the first identifier and the second identifier. This matching algorithm for equipment identifiers updates the digital twin data automatically and continuously by aligning identifiers which refer to the same piece of equipment. The updates flow directly into the digital twin database, thereby removing the manual effort. Using approximate nearest neighbor methods is highly efficient, especially for large plants. The encoder is implemented as an autoencoder which relies only on unlabeled training data. This unsupervised approach is more suitable for industrial scenarios where labeled data is expensive to create.


