Semantic Data Integration Engine for Quality and Interoperability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Enterprise Data Management (EDM) tools struggle with handling incomplete and incorrect data, leading to bad analytics and minimal data interoperability due to bespoke and siloed data models, making it difficult to integrate and aggregate data effectively for both internal insights and external communications.
Innovation Solution
A data integration engine that collects data from multiple sources, curates and links it using semantic mapping, generates unified and interoperable data models, and performs data profiling to address quality issues, enabling de-duplication and providing real-time data management services.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional EDM tools are used to perform data management, then data can be collected and stored, but data quality deteriorates due to incomplete and incorrect data creating feedback loops of bad analytics
Solution Approach 1:
The patent implements feedback mechanisms where data quality metrics are continuously monitored and fed back into the data management system. This allows the system to automatically adjust data collection, validation, and cleaning processes based on observed quality issues, breaking the feedback loop of bad analytics by converting negative feedback into corrective actions that improve data quality over time
Solution Approach 2:
The system performs self-service data quality management through automated profiling, validation, and cleaning operations. The data management system independently identifies and corrects common data quality issues without requiring constant manual intervention, enabling the system to maintain and improve its own data quality while scaling data volume
2Adaptability or versatility
If bespoke and siloed data models are used in EDM tools, then data can be managed within specific business units, but data interoperability deteriorates making integration difficult
Solution Approach 1:
The patent implements a universal data model framework that can accommodate multiple business unit requirements while maintaining interoperability. The system uses standardized data schemas, common ontologies, and unified data structures that work across different domains, allowing the same data model to serve multiple functions and business units without creating silos
Solution Approach 2:
The system introduces intermediary layers including data translation services, semantic mapping mechanisms, and integration middleware that bridge between different data models and standards. These intermediaries enable seamless data exchange and integration between previously siloed systems while preserving the flexibility of individual business unit data models
3Quantity of substance
If external data collection is performed manually by data scientists, then data can be gathered from external sources, but productivity deteriorates due to labor-intensive processes
Solution Approach 1:
The system implements self-service external data collection through automated web crawlers, API integrations, and partner network connections. The data management system independently discovers, collects, and ingests data from external sources without requiring manual data scientist intervention, dramatically improving collection efficiency while maintaining data diversity
Solution Approach 2:
The system performs preliminary actions by pre-configuring data collection pipelines, establishing data sharing agreements in advance, and setting up automated data feeds from external sources. This preparatory work enables continuous automated data collection without requiring ongoing manual effort, improving productivity while maintaining access to diverse external data
Data Source
AI summary
Embodiments herein relate to data management and, more particularly, to collecting data from a plurality of sources, and linking the collected data to derive information and knowledge. The method includes defining at least one data model and asset by including data models, vocabulary, data quality rules, data mapping rules for at least one of, a particular data industry, a data domain, or a data subject area, importing data from a plurality of data sources, performing de-duplication of the imported data and data profiling of the imported data, and creating linked data either by semantic mapping, or by curating the data.


