Semantic Data Integration Using AI-Generated Data Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of integrating disparate datasets across various systems and formats in a cost-effective and efficient manner, without the need for extensive manual data engineering, is not adequately addressed by existing technologies.
Innovation Solution
The use of machine learning and artificial intelligence techniques to automate the generation of semantic models that harmonize and unify datasets, enabling seamless querying and exploration without custom code, by classifying data assets, generating domain-specific models, and mapping equivalencies between datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated machine learning techniques are used to generate semantic models, then data integration speed and efficiency are improved, but the complexity of the system increases
Solution Approach 1:
The patent introduces an intermediary semantic model layer that sits between disparate data sources and end applications. This semantic model acts as a mediator that automatically maps and harmonizes data from multiple sources using machine learning, eliminating the need for complex custom integration code for each application while maintaining high integration speed.
Solution Approach 2:
The system employs self-service mechanisms where the semantic model automatically discovers, maps, and integrates data from diverse sources without requiring manual intervention for each integration scenario. The machine learning components autonomously learn data relationships and generate integration mappings, reducing both manual effort and system complexity.
2Manufacturing precision
If extensive manual data engineering is performed to harmonize datasets, then data integration accuracy is improved, but the time and financial resources required increase
Solution Approach 1:
The patent applies preliminary action by pre-building the semantic model that encapsulates data harmonization logic before actual data integration is needed. The machine learning system performs preliminary learning and mapping during model training, so that when integration is required, the work is already done, achieving both high accuracy and speed without extensive manual engineering at integration time.
Solution Approach 2:
The patent substitutes manual mechanical data engineering processes with automated machine learning systems. Instead of human engineers manually mapping and harmonizing data, the system uses AI algorithms to automatically discover data relationships, perform mapping, and generate integration logic, thereby maintaining accuracy while dramatically reducing time and resource requirements.
3Adaptability or versatility
If custom integration strategies are developed for each application, then adaptability to specific needs is improved, but the overall system complexity and cost increase
Solution Approach 1:
The patent implements universality through the semantic model that serves multiple applications simultaneously. The machine learning-based semantic model learns general data relationships and mappings that can be reused across different applications and use cases, providing adaptability to specific needs through configuration rather than custom development, thereby reducing overall system complexity.
Data Source
AI summary
The systems, methods and computer implemented approaches described herein are directed to tools for data accessibility and exploration that operate by eliminating traditional data engineering delays. In particular, the described approaches accomplish this task without the necessity of generating database specific code. By way of non-limiting example, the systems and methods described are directed to the automated generation of semantic models from structured data using a plurality of machine learning and artificial intelligence technique. It will be appreciated that these machine learning and artificial intelligence tools, when combined, address a unique and difficult problem encountered in the field of database querying and management.


