Adaptive Database Matching for Cross-Nomenclature Object Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In asset life cycle management of large industrial facilities, matching and tracing components across databases with different nomenclatures and standards is a cumbersome, skill-based, and time-consuming process due to varying descriptions and lack of standard conventions for abbreviations and acronyms, making it complex to identify and correlate objects across systems.
Innovation Solution
An adaptive database matching (ADM) system that accesses multiple databases with different schemas, identifies expressions encoded with alphanumerical characters, determines clusters of contextually similar objects, and creates a relational database using a mapping function to establish relationships between these clusters and schemas, facilitating comprehensive data management and integration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If manual matching methods are used to correlate objects across databases with different nomenclatures, then flexibility in handling various standards is maintained, but the process becomes cumbersome, skill-based, and time-consuming
Solution Approach 1:
The patent replaces manual mechanical matching processes with an automated computational system that uses machine learning models, natural language processing, and algorithmic clustering to automatically correlate objects across databases with different nomenclatures and standards, eliminating the need for manual skill-based matching while maintaining adaptability to various standards
Solution Approach 2:
The patent introduces an intermediary adaptive matching system that acts as a mediator between databases with different nomenclatures. This system learns domain-specific languages and standards, translating and correlating objects across different database schemas through learned mappings and contextual understanding, rather than direct manual comparison
2Measurement precision
If comprehensive database schemas with multiple attributes are used to describe objects, then accuracy of object description is improved, but the complexity of matching and tracing components across databases increases
Solution Approach 1:
The patent segments the complex matching process into multiple independent processing stages: expression identification, clustering, mapping function derivation, and correlation. Each stage handles specific aspects of the matching problem separately, reducing overall complexity while maintaining comprehensive object description accuracy through cumulative processing
Solution Approach 2:
The patent transforms the matching problem by changing parameters from direct attribute comparison to learned contextual representations. Machine learning models convert multiple attributes into unified semantic representations, allowing accurate object description matching while simplifying the comparison process through dimensionality reduction and feature extraction
3Productivity
If automated matching systems are implemented, then time consumption is reduced, but the ability to handle domain-specific nomenclatures and standards requires sophisticated algorithms increasing system complexity
Solution Approach 1:
The patent performs preliminary actions by pre-training machine learning models on domain-specific data, pre-establishing clustering structures, and pre-defining mapping relationships. This preliminary processing enables rapid automated matching during operation while the system complexity is concentrated in the initial setup phase rather than ongoing operations
Solution Approach 2:
The patent implements self-service through machine learning models that automatically learn domain-specific nomenclatures and standards from training data without requiring manual configuration. The system adapts to different domains autonomously, reducing the need for complex manual rule-setting while maintaining high matching speed and domain specificity
Data Source
AI summary
A method of associating data from a plurality of databases is disclosed. The method comprises accessing a first database comprising a first dataset and a second database comprising a second dataset. The method further comprises identifying a first set of expressions and a second set of expressions corresponding to the first dataset and the second dataset, respectively. The method further comprises determining a first set of clusters and a second set of clusters corresponding to the first database and the second database, respectively. Furthermore, the method comprises creating a relational database based on a set of relationships and a mapping function determined based on the first set of clusters and the second set of clusters.


