AI Data Matching System for Knowledge Graph Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Complex computing systems face challenges in integrating and aligning data from disparate sources due to issues like incomplete, noisy, and inconsistent real-world data, which requires significant time and expertise for data preparation and mapping, leading to suboptimal use of resources and potential loss of valuable insights.
Innovation Solution
An AI-based data matching and alignment system that preprocesses data, generates feature matrices, and uses techniques like Mahalanobis distance and K Nearest Neighbor to identify similar data sources, building a knowledge graph for structured data integration and enabling downstream applications to gain insights beyond syntactic matching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data preparation and mapping are carried out manually by data experts and data engineers, then data accuracy and domain knowledge utilization are improved, but time consumption and resource requirements increase significantly
Solution Approach 1:
The patent introduces an AI-based intermediary system that acts as a mediator between raw data and downstream applications. This system automatically performs data matching, alignment, and relationship identification using machine learning algorithms, thereby reducing the need for manual expert intervention while maintaining data accuracy. The AI system processes data preparation tasks that traditionally required human experts, significantly reducing time consumption.
Solution Approach 2:
The patent replaces the mechanical manual process of data preparation with an automated AI-based system. Instead of relying on human experts to manually map and align data, the system uses machine learning models, feature extraction, and automated relationship identification to perform these tasks, thereby reducing time consumption while maintaining or improving data accuracy.
2Reliability
If comprehensive data preparation and relationship mapping are performed manually, then data quality and completeness are improved, but expert resource requirements and costs increase
Solution Approach 1:
The patent enables the data preparation system to serve itself by automatically identifying data relationships, extracting features, and aligning data sources without requiring continuous human expert intervention. The AI system autonomously performs data matching, relationship identification, and quality assurance tasks, reducing the complexity associated with manual expert resource requirements while maintaining data quality.
Solution Approach 2:
The patent changes the parameters of the data preparation process by transitioning from manual expert-driven operations to automated AI-based processing. This involves changing the operational mode from human-intensive to machine-intensive, using machine learning algorithms, feature extraction techniques, and automated relationship mapping to maintain data quality while reducing expert resource requirements.
3Ease of operation
If traditional syntactic matching methods are used for data alignment, then processing simplicity is maintained, but insight depth and relationship detection capability are limited
Solution Approach 1:
The patent adds another dimension to data matching by moving beyond traditional syntactic matching to include semantic relationship identification, feature-based alignment, and contextual understanding. The system extracts features from data sources and uses these features to identify relationships that go beyond surface-level syntactic similarities, thereby preventing loss of information and providing deeper insights while maintaining processing simplicity through automated AI-based methods.
4Quantity of substance
If data from multiple disparate sources are integrated without AI-based alignment, then data collection scope is maximized, but data consistency and integration quality deteriorate
Solution Approach 1:
The patent creates a universal AI-based data alignment system that can handle multiple disparate data sources with different formats, structures, and characteristics. The system performs multiple functions including data matching, feature extraction, relationship identification, and consistency validation across diverse data sources, thereby maintaining data collection scope while ensuring data consistency through a unified automated approach.
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
An Artificial Intelligence (AI)-based data matching and alignment system identifies similar data sources for a target data source from a data corpus and generates a knowledge graph that enables downstream applications seamless access to data in the data corpus. The system extracts column features at different levels for the target data source and a plurality of data sources from the data corpus. Feature matrices are built from the features of the target data source and the plurality of data sources. Candidate data sources similar to the target data source are filtered from the plurality of data sources using the feature matrices. The tree-based similarity is estimated and K Nearest Neighbor (KNN) graphs are built to identify columns from the candidate data sources that are similar to columns of the target data source to build the knowledge graph.