Machine-Learned Metadata Clustering for Automated Data Ingestion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data integration systems face challenges in integrating disparate data sources due to varying naming conventions, different data types, and manual, incomplete, and costly integration processes, often requiring wholesale system conversion or significant investment.
Innovation Solution
A metadata processor using unsupervised and supervised machine learning techniques to cluster metadata from different sources, allowing for automated ingestion and integration without disrupting existing systems, while adhering to industry standards.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual integration processes are used to integrate disparate data sources, then some improvement in ease-of-use and accuracy is achieved, but the processes remain expensive, incomplete, and require significant maintenance investment
Solution Approach 1:
The patent replaces manual mechanical integration processes with an automated machine learning-based system. The ML model automatically discovers relationships between data sources, maps schemas, and integrates data without human intervention, eliminating the need for expensive manual processes while maintaining or improving accuracy
Solution Approach 2:
The system enables self-service data integration by allowing the machine learning model to autonomously discover data relationships, generate integration mappings, and adapt to new data sources without requiring manual configuration or expert intervention, thereby reducing maintenance investment
2Reliability
If existing integration reference processes are used to manage integration, then some level of integration is achieved, but it is difficult to implement new integration processes that may be more effective
Solution Approach 1:
The patent implements a dynamic integration system where the machine learning model continuously learns from new data sources and automatically adapts integration strategies. The system can implement new integration processes effectively by training the ML model with additional data, enabling flexible adaptation without disrupting existing reliable integrations
Solution Approach 2:
The machine learning-based integration system provides universal functionality that can handle multiple data sources, schemas, and integration scenarios through a single platform. The ML model learns general patterns that apply across different data types, making the system both reliable for existing integrations and adaptable to new ones
3Productivity
If wholesale system conversion is performed to integrate data sources with different naming conventions and schemas, then complete integration is achieved, but significant investment and disruption are required
Solution Approach 1:
The patent applies preliminary action by training the machine learning model on metadata and schema information before actual data integration occurs. The model pre-discovers relationships and generates integration mappings in advance, enabling complete integration without time-consuming wholesale system conversion or disruption
Solution Approach 2:
The system extracts only the essential integration knowledge from metadata and schema information, rather than requiring complete system conversion. The ML model identifies and extracts key relationships, data mappings, and integration patterns, achieving complete integration while minimizing time investment and disruption
Data Source
AI summary
A system, device and method are provided for assessing actions of authenticated persons within an enterprise system. The illustrative method includes extracting metadata comprising a plurality of categories from a plurality of data sources, and applying an unsupervised machine learning process to the extracted metadata. A plurality of clusters of the plurality of categories of the extracted metadata is generated, and thereafter one or more review criteria are applied thereto to generate curated clusters. The method includes training a supervised machine learning model with the curated clusters. The method includes, in response to receiving a new metadata input, processing the new metadata input with the trained supervised machine learning model. Data associated with the new metadata input is ingested based on respective clusters output by the trained supervised machine learning model for categories of the new metadata input.


